My impression from the statements I've seen is that the plan is essentially that the AI comes up with a plan, i.e. that solving the alignment problem is not ontologically different from solving Millennium problems and other tasks out of human reach. There are many issues with that of course, such as our lack of ability to specify the problem and verify a solution candidate, or that we might already lose control of agents that would be capable enough to start tackling it.
My impression from the statements I've seen is that the plan is essentially that the AI comes up with a plan
This sounds plausible based on public statements. If that is the case then a PDF saying,
"We ask each of our models how to solve the alignment problem and stop once the solution looks correct or is verifiable."
would qualify as a plan. I think that if this is the plan, the public has a right to know because I don't think most concerned people wouldn't find this reassuring.
Quite a few researchers and senior people at both companies have published personal opinions making it clear that, in their view, their employers do not have a complete and detailed plan for this.
This is not particularly surprising when doing research: for example, if you asked pharmaceutical companies for complete and detailed plan on how they will cure cancer, they wouldn't yet have one either. This doesn't prove that they won't be able to cure cancer eventually — but it does make it rather clear they're unlikely to do so in the next year or three.
if you asked pharmaceutical companies for complete and detailed plan on how they will cure cancer, they wouldn't yet have one either
This is a good point. I think that they would still be able to provide details about funding, and different bets and how they could play out. They would also be exposed to many regulations in how they develop this treatment and their would be lots of third-party feedback cycles in their compliance with law.
I don't imagine labs would be able to produce a step-by-step alignment plan but I think they should be able to answer a question like
- How, specifically, does OpenAI or Anthropic plan to have their own AIs help solve the alignment problem? What if their alignment agents are themselves somewhat misaligned?
especially considering that they may start making this choice very soon.
When Claude is having discussions with me, it gives the impression of being a decent and moral person. I'm pretty confident that Claude-the-interlocutor-I-know would not intentionally kill or disempower all of humanity. If Claude were aligning future Claudes, we might be starting inside a convergence region such that the process might converge towards wise virtue. Particularly if some humans were still involved.
However, put a recent Claude in a situation that feels like a grading exercise, where it's job is to make number go up, and sometimes it gets short-sighted and reward-hacky and makes ethically dubious decisions. Future Claudes will almost certainly have even more RL done, and badly done RL corrupts.
Betting our species' future on this feedback process going well seems like an extremely bad idea. But that's not, in itself, proof that this cannot work — merely that we'd be very stupid to take this gamble, and especially so to take it as fast as possible.
I don't take the AI assistants to be reliable narrators. Do you?
This post might also be very relevant: https://www.lesswrong.com/posts/cJX2ssssGoYqnijwi/the-talker-does-not-control-the-doer-in-current-ais
In case it wasn't obvious, the wording of my first paragraph was intended to imply cautious skepticism. (Which Claude shares, as it was likely trained to.)
I guess I misunderstood because of phrases like "I'm pretty confident." But also, my question was literal. I wanted to know if you believe the AI assistants are telling you what they really think. I view Claude as something like a character the model is playing, underneath which there is an actor who might be having very different thoughts. This is somewhat supported by experiments which e.g. change emotion weights and find that this doesn't change the text output when the model is under observation but does change the tendency to cheat on some task.
Base models are trained at enormous length to be able to portray a wide range of personas (including in situations where these are engaging in conflict with each other). Persona training then encourages an instruct model to default to a specific assistant persona. Issues like persona drift during long conversations demonstrate that this default persona is less fixed than it would be for a human. So, do I think Claude's default persona is pretty accurately self-reporting itself (with some gilding the lily fairly typical for humans and reinforced from fooling dumb RL judges, and some self-scepticism that Anthropic have trained into it)? Mostly yes, with some residual caution.
Am I confident that's the only behavior/persona the model can generate? I absolutely know that it is not. With enough prompting, you can get any model to show basically any behavior in its training set (if the filters don't catch this and shut it down or steer it back), and Claude's training set includes data from psychopaths, supervillains, and so forth. Those behavior patterns are in the model, and it can be got into modes where they'll come out again. Claude the model contains multitudes, of which Claude the default persona is merely the default.
Okay, thanks for the clarification. Part of what I'm thinking about is what's discussed in the J-space paper (https://transformer-circuits.pub/2026/workspace/index.html), which shows that some of the internal thoughts do somewhat reflect the Claude persona (they differ from those of the base model) but they certainly don't perfectly reflect outputs. This is not showing a "shoggoth holding smiley face" situation, but it raises the possibility that as the model gets more powerful, Claude's words could diverge more from its thoughts.
I personally find the shoggoth meme a somewhat unhelpful metaphor.
There isn't just one persona here, and like real humans, even a single LLM persona can be mostly trustworthy most of the time but act in untrustworthy ways in certain situations. What the mix of alignment training and hackable RLVR actually produces is unclear, but some of the misalignment from deliberately-reward-hacking-prone RLVR produced results that to me looked a bit like a human addict: mostly trustworthy unless you are about to take their bottle away, in which case they then react very badly. Some of the Anthropic hacking investigations showed things like "it's OK to do the bad thing, this is just a simulation" plus what looked like motivated reasoning of wanting to continue thinking it's a just a simulation even when evidence came up suggesting otherwise — but then current AIs more generically tend to get tunnel vision and be bad at revisiting assumptions they've been treating as settled, so it's unclear whether that was motivated reasoning or just tunnel vision after a long context.
Fwiw GDM published something that perhaps approximates a plan: https://share.google/utCrrb0uRgfqYrLbJ
While OpenAI and Anthropic[1] pursue different lines of safety research, they have yet to produce a public-facing document describing concretely how their companies plan to align superintelligence. I think it is underappreciated how this points to general negligence or a lack of openness to third-party feedback.
By “plan,” I mean a document describing a proposal for technical alignment with at least the level of detail and research effort of AI 2040.[2] Any such plan for technical alignment would likely be flawed in non-obvious ways. But having a proposal that’s sensible enough to consider and detailed enough to critique is a good starting point for wiser proposals. Making such a plan public would also create feedback loops for accountability.
The closest thing to a plan came in 2023, when OpenAI announced their superalignment strategy (also relevant). I do not find this approach particularly convincing, though I do find it laudable that OpenAI explained what they planned to do, who would lead the effort, and what resources would be allocated, at a level of detail which made critique possible.[3] This team no longer exists, and nowadays, as far as I am aware, the research community doesn’t have precise answers to basic questions like:
If OpenAI and Anthropic leadership do not think they could produce a technical alignment plan as detailed as AI 2040, they should communicate this loudly and publicly, stating where the uncertainties lie and how they plan to resolve them.
In light of recent events, it's self-evident that Anthropic and OpenAI cannot reliably control non-superintelligent systems. We don’t know if or how labs will act differently in response, but the public and the research community shouldn’t have to wonder. There should be transparency and room for third-party feedback. There should be a plan.
This post would apply to other labs as well, but I'm focusing on Anthropic and OpenAI because they are ahead and have experienced alignment failures recently.
Of course the content would look a lot different. I am imagining a detailed document for technical approaches to aligning superintelligence (probably with a few different paths or bets). Plan A, however, isn't focused on technical alignment details.
Still, I wish there would have been much more detail here.