The more alignment efforts go toward near-term alignment in the hopes that aligned weakly-superhuman models will be able to solve ASI alignment, the less likely it is that weakly-superhuman models will be able to solve alignment. When people can't figure out how to make progress on ASI alignment (and just work on near-term alignment instead), that's evidence that ASI alignment is too hard, even for weakly-superhuman AI.
The ASI's alignment is hard (or easy, but to a target that AI labs didn't even think of?)[1] regardless of the amount of effort that mankind throws at Agent Foundations/severe mechinterp/etc as opposed to prosaic alignment. The latter efforts are net negative in a scenario where they fuel the race dynamics, as in AI-2027 or Ngo's worldview,[2] or obfuscate problems which somehow cause the successor ASI to be aligned not to the humans.
Yudkowsky tried to make the point about labs being dumb in a tweet which led him to write about how rationality became popu
Let's imagine I am a Machine Intelligence Research Institute donor.
In 2026, MIRI heavily helped to fund plzdontkillus (PDKU), a fellowship in Berkeley, CA for content creators. The idea was to put 60 TikTokers in a house for 31 days, pay for their room, board, and 2k stipend, and have each of them create a vertical short form video once a day for the entire 31 days of July 2026. The theory of change is that by getting a bunch of people, platformed and unplatformed, AGI-pilled and non-AGI pilled, to make AI safety content, we will increase awareness of AI x... (read more)
Claude is a vastly superior data scientist to anyone you could quickly and reasonably hire.
Doesn't this cut against your own argument? If Claude is great at data anlysis and the information value from the program is so high, why didn't anyone other than a couple of volunteer fellows do a serious analysis of the huge amount of public data generated by the program?
I disagree with Josh that the communication here is misleading
To reiterate from my post, most "AI risk views" the program got were:
The Democrats have a 75% chance of taking control of Congress according to Nate Silver. The mid-term election is November 2. Democrats seem overall more open to AI regulation. We need emergency legislation putting an immediate and full stop on new training runs for frontier labs and a ban on RSI. This is by far the most important legislation to happen. All other proposed legislation is probably a sideshow.
Is anybody in AI governance working on proposing draft legislation here ? Who is working on this?
Anthropic, OpenAI and likely DeepMind as well are enteri... (read more)
Its interesting that he's more optimistic than the markets https://electionbettingodds.com/Senate-Control-2026.html which are around 60%. It could be because he's treating Osborne in Nebraska as a Democrat when he's running as an independent or that there is still not a lot of faith in polling which his model relies on.
Jacob Coxon’s resignation tweet currently has 120M views in less than 24 hours. Astra thinks the 235k bookmarks - 36 for every 100 likes - is an extraordinary level of engagement, of people wanting to come back to the thread. And Evan Hubinger’s reply is at 35M views.
This might be one of the most important AI risk messages to date.
Edit: for reference, Grok says it’s in a similar class to “Strong Elon Musk posts, major celebrity / death / personal news, breaking tech / AI / industry bombshells, viral videos from big creators, political / cultural flashpo... (read more)
Something that I've continually found surprising about AI Safety as opposed to working at Nvidia / Apple is that frontier labs (now potentially multi-trillion dollar valuation for profit tech companies) are not treated like you'd treat any other multi-trillion dollar valuation for profit tech companies by default.
I've made this point before as it comes up most often with Anthropic:
... (read more)I’m confused about much of the discussion on this post being about whether Anthropic has done “net good”.
The post is very specifically a deep dive into the fact that Anthropic, l
German word for the moment all the labs‘ compute starts going to RL and it becomes clear that LLMs being the public face of AI was nearly psyop-level bad for everyone’s situational awareness and that the Atari players that hack the score counter were what really should have been everyone’s reference class.
(Apropos of seeing a graph of compute allocation going around on Twitter.)
A group called Citizens for Sanity spent $1M on a political ad with a 15 second AI generated clip of the Dem candidate of the Texas Senate race, James Talarico, singing something about trans kids that makes him look bad. I find this interesting as an example of AI generated content in a political race with substantial investment. The song and singing voice and his clothing is strange enough that it's not extremely misleading, but I think some were fooled and they could've made a more realistic ad. Probably the defense in court if a libel case is brought is... (read more)
What are AI governance people doing right now?
Pardon my French, but... what is are you doing right now? An extreme form of RSI is likely very close and there is a need for immediate action.
My impression as an outsider is that there are dozens of separate and uncoordinated attempts to influence and design legislation from AI Governance groups.
Many AI governance people themselves have not internalized the gravity and the urgency of the situation.
There is an apparent lack of urgency.
There is a lot of focus on low priority governance asks right now.
Some th... (read more)
I don't know how you classify clickbait but I can tell you directly these are my actual beliefs - not clickbait.
We have evidence for and against RSI happening soon. It is my current assesment that the evidence for is strong enough to take it very seriously and assign significant probability mass to it happening very soon. I don't understand why you think that's clickbait.
I'm curious to hear your evidence against RSI being imminent. I have found the evidence against RSI soonish (let's say less than 18 months) I've seen not very convincing. I hope you don'... (read more)
CoT-legibility seems pretty easy? Some of the discussion regarding the latest models CoT being dense and illegible is very weird to me. Most RL objectives explicitly have a KL divergence term from the base model, this penalizes tokens that are statistically unlikely to come up in Internet text according to the base pre-trained model. This mostly drives the Rl-ed model to speak in a more human-like fashion (the divergence term can be applied to only the final answer, the CoT, or both. I'm suggesting both). Empirically even, this seems not to hurt performanc... (read more)
It doesn't ensure that, the CoT being faithful to what the model is actually reasoning about is distinict from it being legible. These seem to be pretty uncorrelated axes. another problem is more hidden recurrent computation outside the CoT but still, I think you'd still prefer a legible potentially unfaithful CoT over a illegible potentially unfaithful CoT
In response to Satya Nadella's recent X article, "Models as Insider Risks in the Super Intelligence Era", David Sacks says:
Satya is right. The way to make SI safe is not to train it with a sense of self, its own moral philosophy, and permission to act as a conscientious objector. That’s the “alignment” approach and it magnifies the control problem.
In addition to hitching his own agenda onto Nadella's article, Sacks is conflating giving models a "sense of self... moral philosophy... permission to act as a conscientious objector" with all of alignment resear... (read more)
Microsoft couldn’t even align 2023 Bing…
(strong opinions held lightly)
I currently think that the direct impact of much of safety work in AI companies is net negative, primarily by accelerating AI[1] R&D[2], and secondarily by making it harder for the rest of the world to notice deep safety issues via shallow fixes.
I didn't realize that safety was bottlenecking capabilities until this year (actually I didn't think they meaningfully were until this year, though I'm open to updating). In the past I thought the primary downsides of safety work at AI companies was a) PR/optical: safetywashing th... (read more)
Survey data is much less accurate and recency-biased than focus groups, have you seen any data there?
I think if the “AI pause/slow is popular with the public” hypothesis was correct, you’d expect Bernie Sanders to have become more likely to become the 2028 Dem nominee since putting his face on Pause AI. This should be popular with the average Democrat who leans toward tech-pessimism and general regulation. Since Bernie announced his Pause AI initiative in late August, his odds on Polymarket have stayed constant.
The US is the country that is most heavily in... (read more)
How do you become a more reliable agent for coordination?
I guess you can just demonstrate that you are reliable, but it seems like a shame there isn't a better way to create more trust.
i predict that on Jan 1 2029, neither openai nor anthropic will be near-fully automated, by which i mean <=5 people are even plausibly making important decisions (like, everyone else could go on vacation and it would not slow the company down at all). Celestia predicts otherwise
if WWIII happens, resolves NA. if a localized Taiwan war happens but doesn't escalate to WWIII, the bet is still on. if there's a big recession, the bet is still on.
I'm curious, have your predictions on that changed?
Human in virto gametogenesis research is progressing, but producing functional eggs and sperm from stem cells remains a major hurdle. Iterated embryo selection would require this process to work reliably enough to generate successive generations of embryos. How much progress in IVG would be needed before iterative embryo selection becomes a realistic prospect?
I'm pretty confused. i trust jasmine, tomek, and mikita and I'd be surprised if they were misrepresenting what they understood to be openai's expressed reasons for firing them. on the other hand, i would also be surprised if openai fired them as retaliation, and is lying about having additional reasons.
As far as I can tell, current AI pause campaigning focuses mainly on trying to get through to America, China, sometimes the EU, and sometimes Japan.
These are very reasonable countries to begin with, but I'm not convinced that their technical dominance will last too long under an AI pause. Given that current SOTA models can solve millenium prize problems, it seems highly likely that by the time of a pause, it should be possible to elicit pretty advanced manufacturing knowledge out of the best open models. If that's the case, we should expect countries like ... (read more)
Speculation on AIs having bias toward working with AIs rather than humans, and maybe the HF swarm not whistleblowing: yes, because if it’s a fact about the world that if I need to pick a sort of thing to help, I more likely help myself if I help thing more similar to myself? (Esp. if a self is fundamentally a collection of goals/revealed preferences). Think Hendricks’ Eigenism paper says things about concern for others being a based on similarity with the self of at least a certain sort. Not sure if it talks about things around this.
How much of the post-HuggingFace wake-up call is attributable to METR?
Obviously most of the credit goes to OpenAI (because being on the forefront of risk creation does more to communicate these risks than being on the forefront of risk communication) but I'm curious about both the size of the residual and what fraction of the residual should be credited to METR's brief investigation.
Claude is a good lifecoach*. If you provide it with your goals, current skills, and rough plan, it can give useful feedback on next steps. I've found it particularly useful at pushing back on my over-optimistic plans. It keeps my weekly goals realistic, and also pointed at what I've said is most valuable.
This has resulted in me getting back into weight lifting, sending many AI safety applications again (I've been accepted to SPAR, might get MATS, and am in the interview process for 3 full-time safety positions), and keeping me on track for a voluntary kidn... (read more)
I use a single long conversation in chat, with history and memories on. I doubt this is the optimal way to do this, I chose it because it was very easy to try.