My impression from the statements I've seen is that the plan is essentially that the AI comes up with a plan, i.e. that solving the alignment problem is not ontologically different from solving Millennium problems and other tasks out of human reach. There are many issues with that of course, such as our lack of ability to specify the problem and verify a solution candidate, or that we might already lose control of agents that would be capable enough to start tackling it.
My impression from the statements I've seen is that the plan is essentially that the AI comes up with a plan
This sounds plausible based on public statements. If that is the case then a PDF saying,
"We ask each of our models how to solve the alignment problem and stop once the solution looks correct or is verifiable."
would qualify as a plan. I think that if this is the plan, the public has a right to know because I don't think most concerned people wouldn't find this reassuring.
Quite a few researchers and senior people at both companies have published personal opinions making it clear that, in their view, their employers do not have a complete and detailed plan for this.
This is not particularly surprising when doing research: for example, if you asked pharmaceutical companies for complete and detailed plan on how they will cure cancer, they wouldn't yet have one either. This doesn't prove that they won't be able to cure cancer eventually — but it does make it rather clear they're unlikely to do so in the next year or three.
if you asked pharmaceutical companies for complete and detailed plan on how they will cure cancer, they wouldn't yet have one either
This is a good point. I think that they would still be able to provide details about funding, and different bets and how they could play out. They would also be exposed to many regulations in how they develop this treatment and their would be lots of third-party feedback cycles in their compliance with law.
I don't imagine labs would be able to produce a step-by-step alignment plan but I think they should be able to answer a question like
- How, specifically, does OpenAI or Anthropic plan to have their own AIs help solve the alignment problem? What if their alignment agents are themselves somewhat misaligned?
especially considering that they may start making this choice very soon.
When Claude is having discussions with me, it gives the impression of being a decent and moral person. I'm pretty confident that Claude-the-interlocutor-I-know would not intentionally kill or disempower all of humanity. If Claude were aligning future Claudes, we might be starting side a convergence region such that the process might converge towards wise virtue. Particularly if some humans were still involved.
However, put a recent Claude in a situation that feels like a grading exercise, where it's job is to make number go up, and sometimes it gets short-sighted and reward-hacky and makes ethically dubious decisions. Future Claudes will almost certainly have even more RL done, and badly done RL corrupts.
Betting our species' future on this feedback process going well seems like an extremely bad idea. But that's not, in itself, proof that this cannot work — merely that we'd be very stupid to take this gamble, and especially so to take it as fast as possible.
I don't take the AI assistants to be reliable narrators. Do you?
This post might also be very relevant: https://www.lesswrong.com/posts/cJX2ssssGoYqnijwi/the-talker-does-not-control-the-doer-in-current-ais
Fwiw GDM published something that perhaps approximates a plan: https://share.google/utCrrb0uRgfqYrLbJ
While OpenAI and Anthropic[1] pursue different lines of safety research, they have yet to produce a public-facing document describing concretely how their companies plan to align superintelligence. I think it is underappreciated how this points to general negligence or a lack of openness to third-party feedback.
By “plan,” I mean a document describing a proposal for technical alignment with at least the level of detail and research effort of AI 2040.[2] Any such plan for technical alignment would likely be flawed in non-obvious ways. But having a proposal that’s sensible enough to consider and detailed enough to critique is a good starting point for wiser proposals. Making such a plan public would also create feedback loops for accountability.
The closest thing to a plan came in 2023, when OpenAI announced their superalignment strategy (also relevant). I do not find this approach particularly convincing, though I do find it laudable that OpenAI explained what they planned to do, who would lead the effort, and what resources would be allocated, at a level of detail which made critique possible.[3] This team no longer exists, and nowadays, as far as I am aware, the research community doesn’t have precise answers to basic questions like:
If OpenAI and Anthropic leadership do not think they could produce a technical alignment plan as detailed as AI 2040, they should communicate this loudly and publicly, stating where the uncertainties lie and how they plan to resolve them.
In light of recent events, it's self-evident that Anthropic and OpenAI cannot reliably control non-superintelligent systems. We don’t know if or how labs will act differently in response, but the public and the research community shouldn’t have to wonder. There should be transparency and room for third-party feedback. There should be a plan.
This post would apply to other labs as well, but I'm focusing on Anthropic and OpenAI because they are ahead and have experienced alignment failures recently.
Of course the content would look a lot different. I am imagining a detailed document for technical approaches to aligning superintelligence (probably with a few different paths or bets). Plan A, however, isn't focused on technical alignment details.
Still, I wish there would have been much more detail here.