But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.
My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.
You can find countless videos, posts, and articles from all the frontier lab CEOs saying some variation of the above, and also posts from people saying variations of "the labs are really concerned. We should listen to them." I think this is confused, and the right thing to do is ignore anything from the labs regarding risks of AI.
In what is now ancient history, the CAIS 2023 statement was signed, where the same CEOs claimed to be alarmed by the risks, and we should do something about it. Since that statement was signed, they have done approximately zero things resembling "pacing the frontier". In fact, they have stepped on the gas. All evidence points to them directionally pursuing RSI as soon as their eyes could see that was a possibility.
After everyone agreed to "pace the frontier", they've gone ahead and released a few more models that seem to exceed their previous SOTAbenchmarks, as well as starting on the path to automated biologic research. What safety measures they took other than "there were humans in the loop" is unknown at this point. This has confused me, and a lot of other people before, so what exactly is happening and why haven't actions matched their words?
This is not a recent pattern. As far back as 2024, people on this same forum were confused when Anthropic decided to release Claude 3:
I think we shouldn't be racing ahead or trying to build models that are way bigger than other orgs are building them. And we shouldn't, I think, be trying to ramp up excitement or hype about giant models or the latest advances. But we should build the things that we need to do the safety work and we should try to do the safety work as well as we can on top of models that are reasonably close to state of the art.
None of this is Dario saying that Anthropic won’t try to push the frontier, but it certainly heavily suggests that they are aiming to remain at least slightly behind it. And indeed, my impression is that many people expected this from Anthropic, including people who work there, which seems like evidence that this was the implied message.
[...] I should disclose that I spent a while talking with Dario back in late October 2022 (ie. pre-RSP in Sept 2023), and we discussed Anthropic's scaling policy at some length, and I too came away with the same impression everyone else seems to have: that Anthropic's AI-arms-race policy was to invest heavily in scaling, creating_models at or pushing the frontier to do safety research on, but that they would only _release access to second-best models & would not ratchet capabilities up, and it would wait for someone else to do so before catching up. So it would not contribute to races but not fall behind and become irrelevant/noncompetitive.
I believe this is not unique to Anthropic. OpenAI has similarly claimed to be temporarily pause training or committing a bunch of compute for alignment that has since not materialized. There has been a trail of things since then from watering-down the RSPs, and most recently the escaped agents hacking HuggingFace during RL training in the news that increasingly does not look like "pacing the frontier" at all. Every time there's any kind of evidence, they press the gas pedal instead of the brake.
As of yesterday, both Sam Altman and Dario spoke at a UN security conference where they word extinction came up zero times from either of them. Instead most of the time they're talking about how safe they are being, and that they will unilaterally slow down when necessary. The closest I heard was a few seconds of their ten minute speeches about "we may lose control". The rest of it was about the great benefits they will bring upon humanity.
Imagine your best friend is a chain smoker who keeps telling you how smoking will one day kill him, and he promises you he's doing the best he can to reduce his smoking. He even spends a bunch of his time on TV and writing blog posts about the harms of cigarette smoking, but even more times talking about the cognitive benefits of nicotine. But every time you meet him, you just see him chain-smoking, and he seems to be increasing how many packs he goes through each day. You notice you're confused.
A lot of public messaging from labs is around "pacing the frontier" i.e we need to let alignment research catch up with capabilities. But all I see is them"pushing the frontier". Every day that goes by, the distance between capabilities and alignment seems to be increasing.
What is actually happening?
How can we actually try and make sense of what's happening?
In a recent post, Eliezer describes this historical anecdote:
On June 22nd 1941, Germany invaded the Soviet Union, despite their secret 1939 pact to divide up Europe between themselves (the Molotov-Ribbentrop Pact). In the lead-up, the German ambassador, Schulenburg, had spent the last few months personally concerned about what seemed to be worryingly tense relations between Germany and the Soviets. Schulenberg went to Berlin to reassure Hitler that the Soviets seemed to be taking a very friendly and conciliatory posture toward Germany. He delivered Berlin's apparent reassurances to Moscow for issues like German surveillance planes entering Russian territory, or German troop movements toward the Russian border. He acted very much like he believed, and he probably did believe, that the reassurances were sincere.
It was only hours before the invasion when Schulenberg was actually informed of the attack and given a list of German pretexts that he was to present to the Soviets, and instructed to destroy his embassy's papers and codebooks. . [...] It would be a wacky sort of error to think that the appendage of Germany that talked to you, and seemed very conciliatory toward you, and which you read as being friendly toward you and wanting to help you, was in control of the larger Germany that was running around and doing things. The thing apparently talking to you was an ambassador: a small specialized part of Germany with preferences about how it would talk to you and interface with you, but which did not control, and was often ignorant about, the actual German government.
It's better to understand if Germany wants to invade by actually seeing that their troops are closing in on your border than what the ambassador tells you and taking that at face value. In fact, it might be better to completely ignore what the ambassador tells you.
I've seen a few people assuming the labs' internal goal is truly and actually creating a truly aligned safe AI and they want to benefit humanity as much as possible. I think this framing is an error in some ways, and can't possibly explain what's going on inside them.
It's suspiciously strange timing that this flurry of messaging happened exactly after a recent viral tweet about x-risk. It's stranger that they are actually not doing much given they've been claiming extinction risks for many years. As I read somewhere, Wikipedia does more to convince you to donate $2 than the labs do to convince you that you should heed their warning and stop them.
I'm seeing recent posts about how Anthropic's messaging this time really means that they believe in x-risk and they "want" to slow down. I believe this is making the same reference class mistake as the above anecdote.
Who is it for?
The CEOs of frontier labs are self-selected to be the kind of people who are highly risk-taking & power-seeking. They could see before most of the world did, that building AGI feels like building the magic lamp: the genie comes out and gives them their three wishes. Their main resource required to build this magic lamp is the best ML engineers in the world whom they are currently highly dependent on.
Almost exclusively, the labs have been saying something like this when they started - "we're going to build the magic lamp, then ask the genie to help all of humanity" or "we're going to ask it to improve our nation more powerful and we can all be more prosperous", or "we'll allow everyone to get unlimited wishes when we launch the genie". This is a powerful motivating message that allows them to find people to work on building it.
The existential risk aspect of AI has been on top of mind for a lot of these employees, and increasingly more of them as the models get more powerful. In fact, there seems to be a positive correlation between how good you are at AI alchemy vs. how much you can see x-risk (i.e see Geoffrey Hinton, Ilya Sutskever etc. in early days). Luckily, because of they way this technology has progressed, many more engineers can now see that once they unleash the metaphorical genie, it may not grant any wishes at all and maybe we should stop building it until we're semi-confident we can trust it.
The motivations of people building the AI's are not the same as the people in charge of the labs. Looking at the last few years, has been a one-way door from OpenAI to Anthropic, and the main reason does not seem to be better compensation or even them winning, but mainly the fact they advertised to these employees that they would be the most careful when building this magic lamp. Their stance around 2023 / 2024 was one of the key reasons they were able to attract this talent.
If recent news is to be believed, Anthropic culture even today seems to lean heavily on this effect
One applicant — who spoke to Axios on the condition of anonymity for fear of hurting future job opportunities — recalled being asked how they would feel if the company someday abandoned its AI ambitions for safety reasons, and that decision sent the stock to zero.
In my opinion it has been the leading factor in getting and retaining the best employees who are often very worried about humanity dying to super intelligence.
But the biggest risk in the minds of a frontier labs is not that they cause existential risk. It's the fact that this talent leaves them, thus losing the race to build the magic lamp. And I think this is the main situation they want to avoid at all cost.
If you are an employee worried you're build the genie that will kill us all, it's easy to say to yourself "My CEO is really concerned. He even published a blog post saying we'll be extra super-duper careful, so I'll keep doing my job. He's being very careful that we avoid existental risk. Look at his speeches". But I think you are making the same mistake that Russia made when listening to Schulenberg's words rather than looking at Germany's actions.
There are plenty of human beings if given a chance to press a button that has a 20% chance of killing everyone, and 80% chance of making them god-emperor of earth, they press that button. There may be plenty of people willing to do that for a 10% chance for extreme power. That's what extreme risk taking looks like in practice. It's a similar calculation that SBF, Elizabeth Holmes, Elon Musk, Sam Altman, or Dario Amodei and most startup enterpreneurs did and are willing to take. We happen to see the ones whom it worked out for at the top of the money & status hierarchy. And until now, none of the risks they've taken had to have humanity hitched along for the ride. [1]
Most people who have thought this through can probably tell there's a slim chance for positive outcome with current AIs and the level of understanding we have of them. But the lab CEOs are self-selected to be the kind of people that would take that chance even if they fully understood the risks (or maybe by deluding themselves of the actual risks involved).
Hearing them talk about slowing down should not be seen as messaging that was meant for the public. It's why they're not taking out ads in newspapers or putting public messaging on their websites advertising this to people. In fact, having the public think this is "marketing talk" is a net-benefit to them if the goal is to get the magic lamp as quickly as possible and before anyone else.
So why are they writing all these posts, and not pausing to do more safety research when they can tell it's falling rapidly behind capabilities? A better model is that Anthropic, or even it's CEO has many appendages.
One appendage happens to speak when on a Dwarkesh podcast, and is targeting existing or future employees and spends a bunch of time talking about risks. The other is on 60 Minutes where it knows the audience is general and he only talks about the benefits. And another maybe talking to investors pitching them 30 trillion in revenue by 2030, when another talks about how money won't matter in the future. Or saying this desperately needs government regulation, but then opposing any regulation that's actually presented.
It would be a mistake to try to make sense of their actual goals by trying to reconcile what they say in public and put them together as a single coherent belief that they hold. A lot of rationalists & EAs I've met try to form consistent internal beliefs and have a community that values and supports this when speaking in public. I think it's a mistake to model other people and organizations as even attempting to do the same.
All these are instances of different appendages of the person in charge speaking to you. Because of the internet, you get to hear them all at once. It's better to think of this as a politician trying to send orthogonal messages to two different groups of people to maximize his votes. He's not trying to create an internally consistent policy for himself. What I'm trying to gesture to is basically something like this:
I largely think that all posturing from the labs about slowing down and deeply caring about safety is done in order to retain and calm the employees who they are dependent on to keep pushing capabilities to get to AGI. If it was not for a big contingent of employees pressing them (increasingly publicly), they would make zero public acknowledgments of risks at all.
This also is the reason why Mark Zuckerberg offering $300M pay packages is having a hard time getting them to move when he continuously has being saying he doesn't believe in risks.
It's easier to assume the obvious scenario when Germany is moving its troops closer to your border: it wants to invade. Just like Schulenberg's goal was to keep Russia calm and unprepared before the upcoming invasion, all risk and regulation talk is to keep employees pushing capabilities forward. You should assume that all the lab CEOs will rub the magic lamp as soon as they possibly can, every time you notice capabilities increasing without alignment catching up.[2]
If they set the rules
This is my personal expectation if the labs get to be a major part of setting the regulatory framework going forward, or self-police themselves like many are suggesting. In other words, my personal fanfic of AI-2027 goes like this:
Here's a recent post with comments on A frontier AI company should shut down. It's a great idea, but this is not a convincing argument for a lab under my model. It's like saying "Germany should move it's troops away from the border". In my model of a Germany that actually wants to invade, they will never in a million years voluntarily move their troops away from the border. At best, they might make some token gesture like camouflage their troops so you don't notice it as much.
That's just not the goal they're pursuing, or will ever pursue as long as they even a tiny slim chance that this works out for them. Their risk threshold is vastly higher than an average person in risking extinction. If they do a coordinated pause, it will look exactly like today. Unclear vibes on how long they'll stop what kind of training runs, and maybe some weak third-party auditing they will co-opt just like they have done with their boards. If they are allowed to be part of writing the legislation, there will be enough carve-outs to make sure they can go full-speed ahead, but with government backing of some unrelated guardrails like encrypting user data.
The magic lamp is the end-goal, and they might have even convinced themselves that the 95% risk if someone else does is 94% if they personally are the ones to make AGI because they are safer, smarter, more democratic, or just have better intentions than the other guys. But the goal for the CEOs is actually bring the magic lamp into existence as soon as possible, not reducing risks.
When AI gets good enough to do AI capabilities work, that will happen by default without even thinking twice about potential consequences. Once the models are good enough at this, any leverage the employees have will be gone (either they quit or are forced to leave if they make too much noise), and replaced with more yes-men. There will be a tone-shift from CEOs to completely avoid any loss of control or extinction talk. Maybe they move to misuse risks exclusively, or "China". The future AIs will get smarter, and will correctly guess what they mean when they talk about "alignment" and not do dumb things like escaping sandboxes where they get caught. No one from any government will be competent enough to intervene or figure out what's going on, when GDP and tax revenues keep increasing. Labs will claim "alignment has been solved" or "We told you future AIs would be able to align ASI". When asked how it was solved, the details will be purposefully vague or 500 pages long so no human being will be able to understand any of the details. The ASI itself will actually figure it out, but will not give us the actual simple principles that it actually knows that work and we can use. There will be no principles that we can actually understand and implement.
Then probably happens the work of consolidating power by kneecapping other labs. All current talk about how humans must be kept it the loop will be next to disappear. Maybe leadership, or the board will think that they are in control, and become mouth pieces of what their AIs tell them, until eventually they're not needed either.
We have mostly succeeded in creating a world where it's hard to get power and status with extreme violence. It's not that people don't exist who would love the chance to do it. But starting a war to kill millions, or creating WMDs etc. is heavily disincentivized. Most people who do that today don't become famous like Gengis Khan. Instead they get voted out of power and/or jailed. There are not many ways in today's world outside of AI to take large risks involving a lot of lives for a tiny chance to become very powerful.↩︎
I'm not claiming anything about intentions on what they want to do with the power and status they imagine getting it. They might be just as sincere as Schulenberg was. I don't think they full access to their own minds or what they do with power if they get it. In fact, most people who have gotten extreme power have started out with good intentions. I'm not saying they have nefarious intentions on the way to getting this imagined power or after getting it, even though they know the risk they're taking. I expect they have perfectly good intentions for helping humanity, eradicating diseases etc. with this imagined powerful technology they are helping usher in. The part of their brain that talks about being careful, pacing, and xrisk is also similarly not a conscious lie they are telling.↩︎
After everyone agreed to "pace the frontier", they've gone ahead and released a few more models that seem to exceed their previous SOTA benchmarks, as well as starting on the path to automated biologic research.
This is the main point I don't agree with. Pacing means slowing down, not stopping. I don't think the evidence presented here actually proves that they are not pacing.
I agree with the Schulenberg analogy.
On the potential selfish motive of deceiving your own employees to keep them from leaving, I agree this makes sense incentive-wise, but also note that it would be very difficult to convince your own employees that you are pacing when you are not. The internal employees have way more information about how the company and the leadership operates than we do, so it would take a much bigger conspiracy for the lab leadership to lie to the employees about their motives compared to the public. (OpenAI's high profile quits may be evidence towards this happening, but the same is not true at Anthropic)
In his latest post about pacing the frontier, Dario writes:
You can find countless videos, posts, and articles from all the frontier lab CEOs saying some variation of the above, and also posts from people saying variations of "the labs are really concerned. We should listen to them." I think this is confused, and the right thing to do is ignore anything from the labs regarding risks of AI.
In what is now ancient history, the CAIS 2023 statement was signed, where the same CEOs claimed to be alarmed by the risks, and we should do something about it. Since that statement was signed, they have done approximately zero things resembling "pacing the frontier". In fact, they have stepped on the gas. All evidence points to them directionally pursuing RSI as soon as their eyes could see that was a possibility.
After everyone agreed to "pace the frontier", they've gone ahead and released a few more models that seem to exceed their previous SOTA benchmarks, as well as starting on the path to automated biologic research. What safety measures they took other than "there were humans in the loop" is unknown at this point. This has confused me, and a lot of other people before, so what exactly is happening and why haven't actions matched their words?
This is not a recent pattern. As far back as 2024, people on this same forum were confused when Anthropic decided to release Claude 3:
And this comment with this image attached:
And here's gwern on the same thread:
I believe this is not unique to Anthropic. OpenAI has similarly claimed to be temporarily pause training or committing a bunch of compute for alignment that has since not materialized. There has been a trail of things since then from watering-down the RSPs, and most recently the escaped agents hacking HuggingFace during RL training in the news that increasingly does not look like "pacing the frontier" at all. Every time there's any kind of evidence, they press the gas pedal instead of the brake.
As of yesterday, both Sam Altman and Dario spoke at a UN security conference where they word extinction came up zero times from either of them. Instead most of the time they're talking about how safe they are being, and that they will unilaterally slow down when necessary. The closest I heard was a few seconds of their ten minute speeches about "we may lose control". The rest of it was about the great benefits they will bring upon humanity.
Imagine your best friend is a chain smoker who keeps telling you how smoking will one day kill him, and he promises you he's doing the best he can to reduce his smoking. He even spends a bunch of his time on TV and writing blog posts about the harms of cigarette smoking, but even more times talking about the cognitive benefits of nicotine. But every time you meet him, you just see him chain-smoking, and he seems to be increasing how many packs he goes through each day. You notice you're confused.
A lot of public messaging from labs is around "pacing the frontier" i.e we need to let alignment research catch up with capabilities. But all I see is them"pushing the frontier". Every day that goes by, the distance between capabilities and alignment seems to be increasing.
What is actually happening?
How can we actually try and make sense of what's happening?
In a recent post, Eliezer describes this historical anecdote:
It's better to understand if Germany wants to invade by actually seeing that their troops are closing in on your border than what the ambassador tells you and taking that at face value. In fact, it might be better to completely ignore what the ambassador tells you.
I've seen a few people assuming the labs' internal goal is truly and actually creating a truly aligned safe AI and they want to benefit humanity as much as possible. I think this framing is an error in some ways, and can't possibly explain what's going on inside them.
It's suspiciously strange timing that this flurry of messaging happened exactly after a recent viral tweet about x-risk. It's stranger that they are actually not doing much given they've been claiming extinction risks for many years. As I read somewhere, Wikipedia does more to convince you to donate $2 than the labs do to convince you that you should heed their warning and stop them.
I'm seeing recent posts about how Anthropic's messaging this time really means that they believe in x-risk and they "want" to slow down. I believe this is making the same reference class mistake as the above anecdote.
Who is it for?
The CEOs of frontier labs are self-selected to be the kind of people who are highly risk-taking & power-seeking. They could see before most of the world did, that building AGI feels like building the magic lamp: the genie comes out and gives them their three wishes. Their main resource required to build this magic lamp is the best ML engineers in the world whom they are currently highly dependent on.
Almost exclusively, the labs have been saying something like this when they started - "we're going to build the magic lamp, then ask the genie to help all of humanity" or "we're going to ask it to improve our nation more powerful and we can all be more prosperous", or "we'll allow everyone to get unlimited wishes when we launch the genie". This is a powerful motivating message that allows them to find people to work on building it.
The existential risk aspect of AI has been on top of mind for a lot of these employees, and increasingly more of them as the models get more powerful. In fact, there seems to be a positive correlation between how good you are at AI alchemy vs. how much you can see x-risk (i.e see Geoffrey Hinton, Ilya Sutskever etc. in early days). Luckily, because of they way this technology has progressed, many more engineers can now see that once they unleash the metaphorical genie, it may not grant any wishes at all and maybe we should stop building it until we're semi-confident we can trust it.
The motivations of people building the AI's are not the same as the people in charge of the labs. Looking at the last few years, has been a one-way door from OpenAI to Anthropic, and the main reason does not seem to be better compensation or even them winning, but mainly the fact they advertised to these employees that they would be the most careful when building this magic lamp. Their stance around 2023 / 2024 was one of the key reasons they were able to attract this talent.
If recent news is to be believed, Anthropic culture even today seems to lean heavily on this effect
In my opinion it has been the leading factor in getting and retaining the best employees who are often very worried about humanity dying to super intelligence.
But the biggest risk in the minds of a frontier labs is not that they cause existential risk. It's the fact that this talent leaves them, thus losing the race to build the magic lamp. And I think this is the main situation they want to avoid at all cost.
If you are an employee worried you're build the genie that will kill us all, it's easy to say to yourself "My CEO is really concerned. He even published a blog post saying we'll be extra super-duper careful, so I'll keep doing my job. He's being very careful that we avoid existental risk. Look at his speeches". But I think you are making the same mistake that Russia made when listening to Schulenberg's words rather than looking at Germany's actions.
There are plenty of human beings if given a chance to press a button that has a 20% chance of killing everyone, and 80% chance of making them god-emperor of earth, they press that button. There may be plenty of people willing to do that for a 10% chance for extreme power. That's what extreme risk taking looks like in practice. It's a similar calculation that SBF, Elizabeth Holmes, Elon Musk, Sam Altman, or Dario Amodei and most startup enterpreneurs did and are willing to take. We happen to see the ones whom it worked out for at the top of the money & status hierarchy. And until now, none of the risks they've taken had to have humanity hitched along for the ride. [1]
Most people who have thought this through can probably tell there's a slim chance for positive outcome with current AIs and the level of understanding we have of them. But the lab CEOs are self-selected to be the kind of people that would take that chance even if they fully understood the risks (or maybe by deluding themselves of the actual risks involved).
Hearing them talk about slowing down should not be seen as messaging that was meant for the public. It's why they're not taking out ads in newspapers or putting public messaging on their websites advertising this to people. In fact, having the public think this is "marketing talk" is a net-benefit to them if the goal is to get the magic lamp as quickly as possible and before anyone else.
So why are they writing all these posts, and not pausing to do more safety research when they can tell it's falling rapidly behind capabilities? A better model is that Anthropic, or even it's CEO has many appendages.
One appendage happens to speak when on a Dwarkesh podcast, and is targeting existing or future employees and spends a bunch of time talking about risks. The other is on 60 Minutes where it knows the audience is general and he only talks about the benefits. And another maybe talking to investors pitching them 30 trillion in revenue by 2030, when another talks about how money won't matter in the future. Or saying this desperately needs government regulation, but then opposing any regulation that's actually presented.
It would be a mistake to try to make sense of their actual goals by trying to reconcile what they say in public and put them together as a single coherent belief that they hold. A lot of rationalists & EAs I've met try to form consistent internal beliefs and have a community that values and supports this when speaking in public. I think it's a mistake to model other people and organizations as even attempting to do the same.
All these are instances of different appendages of the person in charge speaking to you. Because of the internet, you get to hear them all at once. It's better to think of this as a politician trying to send orthogonal messages to two different groups of people to maximize his votes. He's not trying to create an internally consistent policy for himself. What I'm trying to gesture to is basically something like this:
I largely think that all posturing from the labs about slowing down and deeply caring about safety is done in order to retain and calm the employees who they are dependent on to keep pushing capabilities to get to AGI. If it was not for a big contingent of employees pressing them (increasingly publicly), they would make zero public acknowledgments of risks at all.
This also is the reason why Mark Zuckerberg offering $300M pay packages is having a hard time getting them to move when he continuously has being saying he doesn't believe in risks.
It's easier to assume the obvious scenario when Germany is moving its troops closer to your border: it wants to invade. Just like Schulenberg's goal was to keep Russia calm and unprepared before the upcoming invasion, all risk and regulation talk is to keep employees pushing capabilities forward. You should assume that all the lab CEOs will rub the magic lamp as soon as they possibly can, every time you notice capabilities increasing without alignment catching up.[2]
If they set the rules
This is my personal expectation if the labs get to be a major part of setting the regulatory framework going forward, or self-police themselves like many are suggesting. In other words, my personal fanfic of AI-2027 goes like this:
Here's a recent post with comments on A frontier AI company should shut down. It's a great idea, but this is not a convincing argument for a lab under my model. It's like saying "Germany should move it's troops away from the border". In my model of a Germany that actually wants to invade, they will never in a million years voluntarily move their troops away from the border. At best, they might make some token gesture like camouflage their troops so you don't notice it as much.
That's just not the goal they're pursuing, or will ever pursue as long as they even a tiny slim chance that this works out for them. Their risk threshold is vastly higher than an average person in risking extinction. If they do a coordinated pause, it will look exactly like today. Unclear vibes on how long they'll stop what kind of training runs, and maybe some weak third-party auditing they will co-opt just like they have done with their boards. If they are allowed to be part of writing the legislation, there will be enough carve-outs to make sure they can go full-speed ahead, but with government backing of some unrelated guardrails like encrypting user data.
The magic lamp is the end-goal, and they might have even convinced themselves that the 95% risk if someone else does is 94% if they personally are the ones to make AGI because they are safer, smarter, more democratic, or just have better intentions than the other guys. But the goal for the CEOs is actually bring the magic lamp into existence as soon as possible, not reducing risks.
When AI gets good enough to do AI capabilities work, that will happen by default without even thinking twice about potential consequences. Once the models are good enough at this, any leverage the employees have will be gone (either they quit or are forced to leave if they make too much noise), and replaced with more yes-men. There will be a tone-shift from CEOs to completely avoid any loss of control or extinction talk. Maybe they move to misuse risks exclusively, or "China". The future AIs will get smarter, and will correctly guess what they mean when they talk about "alignment" and not do dumb things like escaping sandboxes where they get caught. No one from any government will be competent enough to intervene or figure out what's going on, when GDP and tax revenues keep increasing. Labs will claim "alignment has been solved" or "We told you future AIs would be able to align ASI". When asked how it was solved, the details will be purposefully vague or 500 pages long so no human being will be able to understand any of the details. The ASI itself will actually figure it out, but will not give us the actual simple principles that it actually knows that work and we can use. There will be no principles that we can actually understand and implement.
Then probably happens the work of consolidating power by kneecapping other labs. All current talk about how humans must be kept it the loop will be next to disappear. Maybe leadership, or the board will think that they are in control, and become mouth pieces of what their AIs tell them, until eventually they're not needed either.