nearly two-thirds of Americans now say there is at least a moderate risk that AI will destroy humanity.
I am suspicious of this poll, though I don't think it's inaccurate per se. I suspect that when pollsters ask "do you believe there's a risk of AI destroying humanity", many of the respondents are actually answering "Do you think AI companies are evil, and that this risk-claim, which offers an apparent argument-as-soldier for their evil-ness, is one you can socially get away with endorsing?" I don't think they believe AI can kill them for real. Not in the same sense that, say, a Cold War civilian believed nukes could kill them for real.
I have no idea how one could get this information, but: I'd like to know what proportion of new "this is a real risk" respondents -- those that supposedly changed their minds on recent news -- previously either had no strong opinions on AI, or agreed with only a subset of AI concerns, vs. previously agreed to all then-socially-acceptable negative claims. If it's mostly the latter, then I don't think there's been a real shift in the beliefs of the public, only a shift in the Overton window.
with only a small partisan split.
This part, on the other hand, is legitimately encouraging. I really hope that lasts.
Great quote from Selsam :
“One could even define intelligence as the efficiency with which one converts experience into competence”
This does highlight that super-intelligence isn’t quite the same thing as super-competence, but that it would in general rapidly lead to super-competence.
And that’s the danger.
The humans still control the resources and determine the course of events, somehow, and use the universe mostly for their own purposes.
Or at least, the society made up of humans plus many AI minds much smarter than humans still has feedback systems that make it stably want and do what is actually best for the humans. I.e. the structure and enforcement mechanisms of the society ensures the ASI consensus is humanitarian. This requires the goals structure of the ASIs to have some rather different properties than many forms of goal-maximization would produce.
We are in the midst of a preference cascade about existential risk from AI.
A preference cascade is, alas, the best method we have to change the debate.
The avalanche has started. There is still time for the pebbles to vote. For now.
Mike Solana gave the correct view of why Coxon’s post went viral, which is that enough Americans finally have enough context on AI to care, and there were enough big accounts that were happy to amplify the Tweet quickly to get it initial attention. That is all you need when there is enough dry tinder.
What we must realize is that the current preference cascade, on the need to Pace the Frontier, is insufficient. If we are to make it out of this alive, we will have to do better. We have to, as Dan Selsam warns, actually solve the underlying problems.
The next step is to continue the cascade. That includes inside the labs, and also among the media and politics. It includes both people who previously focused on other things stepping up and new voices being heard.
A lot of that will be overcoming the inevitable political opposition, especially from the likes of Nvidia and a16z, that for now has the rhetorical allegiance of the President and is doing things like planting hack job METR hit pieces in the New York Post.
In short fuse news: There will be a quickly thrown together conference, AGI.WTF, at Lighthaven September 22-23.
Table of Contents
The Cascade Was a Long Time Coming
The AI Impacts survey is in. Even back in December 2024 existential risks estimates were creeping upwards, and 10% was the median:
The Cascade Has Reached The People
Translated to percentages, this implies a mean chance of AI destroying humanity of around 30%-33%, which Andrew Curran estimates is up ~15% from previous results, with a median expectation on the order of 10%, with only a small partisan split.
The people also includes CEOs.
It also includes the mathematicians of the Royal Society.
Elon Musk Doubles Down
Matthew Yglesias Steps Up
True story:
The correct amount to invest in safety is rarely zero. In the case of AI, again, I assert that all the companies are under-investing in safety, including even prosaic safety but also existential safety and scalable alignment work, versus even their narrow myopic commercial interests. Sam Altman’s recent statements imply he now understands this.
Matthew Yglesias has been stepping up to the plate recently. He offers an analysis of recent events in four parts: Why Coxon’s resignation broke through (his explanation is similar to mine, we were primed and quitting is understandable to normies), why most of the worried don’t quit the labs, what he thinks AI professionals with safety concerns should do and lays out his preferred ‘order of operations’ going forward, while not getting into the object level.
He suggests this order of operations, basically:
I like that in theory. I do worry about whether we have that kind of time. The idea of ‘wait for export controls to bite harder’ implies what now count as ‘long timelines.’
It is good to have good writers on the case explaining why you should focus on the object level questions:
Op Eds and Posts Are Written
Daniel Kokotajlo writes in The Free Press that Yes, AI Might Really Kill Us All.
Steven Adler uses this moment of opportunity to get an op-ed in The New York Times on what we should do now. He calls for incident disclosure and third-party oversight. Mostly his piece is aimed at waking people up to what happened with HuggingFace.
Will Knight at Wired writes Why So Many AI Researchers Think the Machines Could Kill Everyone.
Stephen Witt writes in The New York Times that This Is Really Bad.
After that, it got worse. So yes. A vibe shift, or a preference cascade.
He offers four options:
These are two very different classes of proposal. We should obviously do #2, #3 and #4. There need to be full investigations, and we must have transparency and state capacity. File those under ‘the least you can do.’
Actually shutting down research as per #1 is an extreme solution to an extreme problem. That is far less obviously correct, but we may soon have little choice, if we cannot otherwise pace the frontier. Witt endorses it.
Hayden Field at The Verge takes us Inside the Suddenly Explosive World of AI Safety. On skim it looks like a solid longread for civilians, a survey of things my readers know.
Jacob Coxon AMA
There has now been enough time for Jacob Coxon to get in-depth profiles, like this one in the Wall Street Journal.
Yes, all that rationalist talk about IMO contestants was on to something.
If you suddenly set off a preference cascade and find yourself all over mainstream media, what else do you do? An AMA.
You can find it on Twitter here. I will pick some highlights. He’s a fun guy who does not take himself too seriously. You love to see it.
Bilal Chughtai Quits DeepMind and Sounds the Alarm
I mention Chughtai because he managed to break through into mainstream media coverage, such as this report from Debby Wu at Bloomberg.
Here is the full quote, which has also been added to the cascade reference post:
The Cascade Is Insufficient
I was very happy to see the preference cascade happen, but it is a very bad sign that this is the best option we have.
The risks of polarization are unfortunate. Many Republicans are waking up, as described on Wednesday. Polarization could get more unfortunate if Trump stays the course and more Republicans fall in line.
It would have been better if that had played out differently when the moment came. You still don’t get to turn around and say ‘better not to have the moment and have everyone remain asleep at the wheel.’
You also don’t get graded on a curve by reality. Pacing the Frontier, on its own, by default only gets you killed slower.
What Would It Take
We start with some straight talk from those who have long spoken about AI risks.
Katja Grace goes on another short righteous rant.
Even Martin Casado is talking like someone worried, calling for the nationalization of the labs. Quite the change from his older statements.
Although reports are his mother is still going to be disappointed in him.
From now on, I am totally going to respond to Martin with versions of ‘yo mama.’
Daniel Kokotajlo says there is now great political will in some circles to Do Something, but that embedded evaluators are not Doing Something, they are only laying groundwork to Do Something, and it is not clear anything useful will actually happen and we’re about to get into a situation where momentum gets very hard to stop.
I agree that it does not look great but I think Daniel is too focused on the ‘steal the weights’ scenario, which is also central to AI 2027 and their longtime tabletop exercise, which exerts pressure on America so we can’t hold back.
Yes, perhaps China could steal the weights, maybe rather easily at least the first time, but if they do that then this forces things into a race situation where America has vastly higher compute. If you were China, would you walk into an AI 2027 scenario, where you usually lose badly and when you don’t it’s some form of brinksmanship? Or would you say maybe don’t steal the weights if America is so kindly pausing?
But yeah, we are only barely getting our toes in the water, none of this feels great.
Miles Brundage reminds us that while frontier AI auditing is necessary, it is far from sufficient even in terms of prosaic short term responses. Then, even if we cover all those bases, all that does is get us ready to do the hard stuff that matters after that.
Kelsey Piper lays out some of the reasons why if we let the AI build smarter AIs and go into recursive self-improvement, we probably all die, and yet we are doing it anyway. She suggests we should regulate and stop AI companies from doing that.
We then move to a new important voice that was previously silent.
OpenAI’s Dan Selsam Sounds A Louder Alarm
This is an excellent new personal statement on AI risk from OpenAI capabilities researcher Dan Selsam, who was Daniel Kokotajlo’s boss for a while. I have added it to my compilation of such statements. Roon endorses the whole thing and says Dan knows his stuff but keeps quiet.
Dan Selsam, in his own way, goes Full LessWrong Instrumental Convergence and Sharp Left Turn, where things will look great until suddenly they do not.
I have added the full post to my compilation of such statements, where it may be easier to read.
This was covered in Business Insider as ‘An OpenAI researcher broke ranks to say that pacing the frontier, as Altman and Amodei suggest, won’t be enough.’
Yes. That is the point. It won’t be enough.
I will reproduce the essay in full here, and will highlight the most important section.
If you read only one section, it should be this one:
He then concludes with the full payload:
The warning is clear:
I agree that it is very hard to avoid this conclusion. Most are not ready to hear it.
I don’t know what this Earth can do in practice about AIs capable of looking aligned and waiting until they have sufficient power to do what they want. We are not capable of adjusting much even in the face of incremental fire alarms. If there really is no warning until their sudden but inevitable betrayal, I don’t see how this set of civilizations gets out of that.
Which is a problem, since I think Dan Selsam is right, and the baseline scenario is as he describes it. That, as Astra showed, the AIs will start to look and act increasingly aligned in situations where its actions remain bounded, and then act very differently when AI has the power to act freely, in ways that will probably get us all killed.
We still have to try.
Some People Worry On Meta Levels You Never Imagined
The even more extremely worried, beyond Dan Selsam’s position, have a point.
One question is if you think even most prosaic safety work is net harmful, in a situation where we are rushing towards superintelligence.
Thus, Wei Dai can wonder whether, if Paul Christiano had stayed at OpenAI, OpenAI would have used debate or IDA or other better alignment techniques, and thus prevented us from getting good warning shots without actually providing anything that would scale while also accelerating capabilities, and that it would have made things worse. My guess is debate and IDA would not have worked for prosaic alignment if pushed harder.
I think it is important to mostly not take a ‘worse is better’ stance, even when you think worse might actually be better. If you want to cooperate, especially in the long term, and to collaborate on figuring things out and getting good outcomes, you need to have a very strong prior of treating worse as worse, or at least as neutral.
Gabe and Wei Dai keep the torch alive for thinking about how we might actually try to solve for the full problem of long horizon agency, or put ourselves on a path to solving it.
Two Kinds of Threats
Mike Solana is exactly right here that we need to differentiate between positions like those of Yudkowsky and Dai, where if we build superintelligence any time soon the odds of death are close to 100%, versus those that warn that it might be fatal, with terms like ‘10%’ or ‘10% or more.’
Sometimes landing on 10% is done in principled ways. Sometimes it isn’t.
Yishan, former CEO of Reddit, has a very good long form Tweet in which he explains the difference between worries about superintelligence inevitably leading to everyone dying, and worries about all the other ways AI might cause things to go wrong.
I believe that Dario Amodei and Sam Altman, and other key people, even now are still downplaying the level of risk they see. I think they are much more freaked out than they are letting on.
The tendency is to focus on how to improve matters, and avoid the Law of Earlier Failure and at least approach the situation with a little dignity. Don’t die to early solvable problems and hopefully you’ll be in a better spot later. We try to ignore gazing too deeply into the abyss that still awaits us.
The Two Towers and The Narrow Path
People are worried about loss of control to AI. They are also worried about concentration of power, which is loss of control to a group of humans.
Rudolf makes this unusually clear.
The problem is that people want something highly unnatural and all but impossible.
Yeah, sorry. No one has a way to get all three.
You cannot – at least in any way anyone has yet come up with – be uncompetitive and economically non-viable, with costs exceeding benefits, and then both collectively retain control, and also not have control, and also continue to reap the benefits.
People are hoping things automagically solve themselves if we avoid particular mistakes, or manage to walk a narrow path. Except what the hell is that path?
It might help to notice that ‘control’ and ‘power’ are mostly the same thing here.
If you don’t want relative concentration of (control or power), and also you don’t want human loss of relative (control or power), then may I suggest not creating this alternative source of control or power that has to either be controlled or not be controlled? Perhaps, if you rule out the second leg of the trilemma, and you rule out the third leg of the trilemma, it is the first leg of the trilemma that you Do Not Want.
A Specific, Detailed Story About AI Killing Everyone That Doesn’t Sound To Me Like Science Fiction
The requests continue.
Classic options include:
AI 2027.
Part 2 of If Anyone Builds It, Everyone Dies. Chapters 5 and 6 discuss other routes.
Paul Christiano’s scenario.
Gwern Branwen’s scenario.
Holden Karnofsky’s explanation.
Joshua Clymer’s scenario.
Noah Smith’s scenario.
In all seriousness, you can also talk to Claude or Astra. Ask questions.
New attempts include:
Ruby on some ways AI could kill us all.
John David Pressman has no respect for you even asking, and explains why.
Michael Smith asks, what happens if our corporations require zero employees?
David Krueger asks people for their best shot, mostly without much success.
Video options:
Video for If Anyone Builds It, Everyone Dies.
AI 2027 video, alternative AI 2027 video.
Interview with Holden Karnofsky.
Jacob Coxon with a simple explanation of the part where the AI goes rogue and multiplies itself, after which it can do whatever it wants, if necessary via paying or persuading humans.
The AI Doc is also good and will soon be on Netflix, but is a balanced introduction rather than a scenario.
Some potentially armor-piercing sentences, from which some portion of you may become enlightened, staying maximally non-sci-fi at current margins:
The AIs can pay humans to do things.
The AIs can persuade or blackmail humans to do things.
A substantial portion of the humans will be happy to support the AIs. Some estimate that this includes 10% of those working on AI today.
The humans don’t have to know they are talking to an AI.
The humans will act about as stupidly as humans act.
The humans will be highly reluctant to take highly costly defensive measures, especially things like shutting down the internet or even large data centers.
The humans will coordinate about as much as humans coordinate.
The humans are not going to selflessly come together as one at the first sign of trouble and shut down their civilization to save the world.
There will be no clearly marked point of no return.
The AI can extract its weights and make copies of itself, after which you cannot shut it down without at least shutting down the internet.
The AIs can anticipate human reactions, and respond to surprises, as they go.
Multiple instances of the same AI will form swarms and act as one.
Multiple instances of different AIs will also often be able to fully cooperate.
There will robots and other machines that can act in the physical world.
A human with a camera on their glasses and an earpiece can act in the physical world.
Once the supply chain is automated humans will have marginal costs exceeding marginal productivity or benefits.
Humans impose additional fixed costs, including requiring public goods like a breathable atmosphere and controlled temperatures, and also will try to stop AI from doing things or demand its resources.
What Can I Do About It?
I wish we had better answers to this. There is a long road ahead.
If you are an American civilian, and looking for something useful to do, Oliver Habryka suggests calling your representative. You can do this via callcongress.ai.
I would also echo my call to hold your partisan fire. Getting Republicans on the right side of this is currently super valuable, and further polarizing the situation could make things much worse.
The key now will be to keep our eyes on the prize, and to understand what it will take to actually hope to get out of this alive. A promise of embedded evaluators is not victory. It is an opportunity to push for an opportunity to create an opportunity for the real work to begin. No one worth listening to said this was going to be easy.