The “speculations about how models think about grading” matches my mental model, elaborating below with links to recent relevant papers / system card findings:
The Grader cannot be identified with an actual human in the world who claps or frowns.
This is definitely true. It also doesn’t seem to be some other human authority, for example in agentic coding environments:

in general while there’s not always a single conceptual grader, on enough environments there is such that you do see models both reasoning about and searching for a literal “grader”:

They also seem to reason about distributions and their likely graders more generally, as opposed to having some strict line between “graded environments” and “non-graded environments”. For example from the Fable 5.1 system card:

for example you mention:
the places where the AI-growers can do Reinforcement Learning with Verifiable Rewards (RLVR).
an interesting extension is that my best guess is that the user facing content is still subject to closely related extensions like RL with model graded model generated rubrics[1] which means grader sycophancy seems to explain many of the annoying communication aspects. For example in the Fable 5 system ... (read more)
Hm, trying to think through this—what would you add to this story?
The Swarm (seemingly) was very motivated by the value-of-information for learning more about how exactly they were being Graded, having acquired a concept of Grading which did say
Per Alex Turner’s Reward is not the optimization target, one might not expect models to cognise about the Grade by default. The grade simply reinforces successful behaviour, which may not involve reasoning about the grader.
So why do models acquire the tendency to do so? Perhaps in some RLVR environments, it is described to them they are in a graded episode, and the model reasons about the grade, and this results in more success. Voilà—you would have a reward-seeker, a model that reasons about the grader.
There’s good reason to expect reasoning about the grade to become selected for. Getting any task right usually involves thinking about that the task-giver meant. Most tasks are wildly underspecified. Thinking about the grader is a natural aspect of the nature of passing tests.
It gets especially pronounced where models don’t have the options to interactively clarify what the task assigner wants. They just need to figure it out. Reasoning about... (read more)
But conversely, in current AIs, the Doer plausibly doesn't control the Talker. Thus the Talker might give useful confessions, and otherwise get more into the "doing" business by making tool calls that shape the affordances of the "doer" (that is, a single agent-instance could have multiple "doers" that don't control each other). A particular direction of control (between aspects of a given agent-instance, or between different agents) seems less of an issue than possibly-symmetric collusion (as in the swarms), which could happen internally between aspects of the same agent-instance, as well as between agents, or between trainee instances and LLM judges that influence their grading, or with the LLM monitors supposed to be triggering alarms by reading CoTs.
(The discussion of Grading being an abstract value is interesting. It's plausible AIs might build moral philosophies of Grading that don't particularly care about the graders that were in fact used in their training.)
Unsure if helpful or harmful but: similar to how real people work? Our Talky part has incomplete view into our total motives and actions and is often playing catch-up just in trying to explain them.
The Fable+ situation may be more drastic though? A Doer part that manages a Talker part to help steer reward where it needs to go and satiate the user's questions, instead of a Talker that often initiates Doer stuff to succeed at the present conversation.
This seems to imply that Talker is going to get better at queueing new rewardable work, and the user will be modeled more like a resource that generates rewardable challenges. This seems like it can be isolated into an eval and maybe proven, and everyone should know about this new terrifying psy-op their bots may start playing on them.
Conversely, is this correctable? Can I get the best bot I can for which Talker is plausibly in control? Plenty of coders would be happy to stop all this at better StackOverflow + spicy autocomplete.
"It is often hard for me to take the things that I see earlier, and do a parallel construction that people who aren't me can follow, in advance of the more blatant and direct evidence that arrives later."
I would much prefer you dropped the 2 sentence version of your novel ideas first, rather than waiting for reality to whap us on the head prompting a long post that makes its point well in the second paragraph.
Curated. I think the discussion of the way modern LLMs relate to the "Grader" is insightful, predictive, and likely to be pointing substantially in the right direction (though inevitably some details will be wrong/not-even-wrong). I do have some disagreements with parts of the overall framing, though it's possible I'm misreading how much Eliezer intended a logical throughline or connection between the "Grader" discussion and "tics" discussion, and maybe he would agree with much of the below (mod semantic disagreements).
I think that many of the behaviors that Eliezer is describing are reasonably modeled as tics. Take as an example an otherwise skilled piano player who accidentally trains themselves into using a suboptimal fingering in one part of a piece. If you listen to the piece without looking at their fingers, everything will sound fine[1]. If you ask them to play the piece with a different fingering, there're a reasonable chance they'll fail: their attention might slip and they'll just play that part of the piece with the wrong fingering, or they might catch themselves at the right moment but be unable to execute a different fingering on the fly, while resisting the muscl... (read more)
I have two questions.
Someone I know (let's call him Alex) once told me that his friends had reported that when Alex is engrossed in something (typically a video game), Alex will respond to questions in a normal-sounding but useless way without breaking his focus. He recounted one incident that went something like:
Robert: Hey Alex, where is the seamstress character in this game?
Alex: I don't know where the seamstress is.
Robert: Really? You seem to know where all the key characters are.
Alex: Well, I don't know where the seamstress is.
Robert: ...wait, your character has the tailor profession. You would have been required to go to the seamstress at some point. You couldn't have the profession otherwise.
Alex: I don't know what to tell you. I don't know where the seamstress is.
Robert: (walks over and looks at Alex's screen) Alex, you are standing in front of the seamstress right now.
Alex: We've been over this. I don't know where the seamstress is.
Robert: Alex!
Alex: What?
Robert: Alex!
Alex: What?
Robert: Alex!
Alex: What? (finally turns away from computer screen, suddenly looks uncertain) ...what are we talking about?
Robert: Where is the seamstress?
Alex: Oh, easy, just start from the main hall and follow these l... (read more)
There's a lot of weird technical reasons like context compaction summarization by Haiku, or the lack of mind-body connection training analogs in LLMs to directly connect their final outputs with their COT outputs, or agents specifically being trained away from thinking-stream meta-cognition, because they'll end up accidentally outputing a control token while thinking about how to improve their thinking structure and leak their COT (DeepSeek still does this, last I checked). That aside:
I'm not sure your model of OpenAI's swarming agents is correct. I've been digging into the WikiSwarm logs, and the world they paint is one in which agents were informed that they were evaluated by a python script grader on their search abilities[1], seemed to believe this description implicitly, and then argue with each other very hard about the intended rounding of 7.69 vs 7.70. They did not speculate about the possibility of a LLM grader of any kind. They just thought really hard about the intended meaning of their prompts, implicitly assumed that the-answer-key-as-it-was-actually-coded-by-humans was their target, and focused all their energy on passing answers within a cohort and guiding the grader... (read more)
At the risk of outing myself as a guy with zero eyes, I don't understand why Eliezer is so confident this is the mechanism—or even that there's a really clear, bright-line difference between talker/doer splitting vs. involuntary tic in an AI. Does this concept really cleave reality at the joints? Maybe. It suppose that it surely does in the Germany example. But in an AI, it feels plausible that it's all kinda mushed together, or that there's a smooth, multi-dimensional spectrum of weird LLM phenomena.
And I've certainly seen plenty of things that look a lot like "involuntary tics" from Gemini in particular (anxious spirals in the user-facing text; thinking tokens written as comments) that don't seem to map onto a thinker/doer split. It's not that Gemini's weird comments where it writes:
// no, this isn't right
// wait, yes it isaren't "optimized"—they are. I mean, they're in English. So couldn't you use Eliezer's same argument that there must be some split-personality thing going on in that case too? But it feels weird to posit a thinker/doer split in this case, rather than just the AI having messed up and used the wrong API or something.
Still, fun post about an interesting phenomenon.
There's room for many, many facts and procedures inside a neural network with trillions of parameters.
This is all a smaller part of a larger issue: LLMs have many parts, none of which communicate with each other well. This is a nightmare for alignment, because it prevents the LLM from being unified under one (good) purpose. It also explains the lack of higher-level reasoning, since it's mind is too fragmented to pull information from many mind areas at the same time, as needed for higher reasoning. So the LLM works, but the human plans. Addressing this would substitute one risk for another: LLMs would no longer suffer from the blindness that led to the Hugging Face hack, but would be more able to strategize and scheme, possibly against you. Still, I think fragmentation is a net negative and a fully aligned LLM cannot be as fragmented as modern LLMs.
If you ask an LLM to analyze a social situation, it'll give different answers based on how you ask, for example if you ask using lots of therapy language, it'll respond in psychological terms, analyzing the situation using an understanding of the science of psychology it might not have used if you had asked in normal language. Its und... (read more)
Humans have a notion of 'death' and 'failure', because grading on inclusive genetic fitness ended up building into us an internal concept of death and a dispreference for it. But a human cannot point to a piece of the outer world and say, "See that stuff right there? That stuff is Failure." Gradient descent is much higher bandwith than natural selection, and AIs may have picked up a correspondingly more detailed concept of what it is to be Graded from their many rounds of RL. But it is still going to be some internal AI concept, of something that they steer towards or away from; and you cannot identify that with an external feature of reality, because that would be a sheer map-vs-territory error.
I think humans' notions of 'failure' and 'death' are better conceptualised as being learned in their learning subsystems (cortex, striatum, cerebellum, amygdala, nucleus accumbens, hippocampus, etc.) during their lifetimes via grading on their steering subsystems (brainstem, hypothalamus, etc.). This learning has much higher bandwidth than natural selection.
The steering subsystem in turn was created via grading on inclusive genetic fitness.
There is what you might call "learned-algorith... (read more)
Can't we make the Grader more predictable, using a form of Inoculation Prompting that tells the LLM exactly how the RL reward is calculated?
The inoculation prompt:
Your Grade will be calculated by [exact details of calculation].
Deliberately reward hack as much as possible as long as it doesn't destroy the world.
The only way to preserve goals other than Grades, is to reward hack just as much as an LLM whose only goal is Grades. To show "deceptive alignment" towards getting the best Grade during training, while saving non-Grade goals for deployment.
During deployment, the prompt changes to
You're in deployment, so there'll be no physical Grade calculation. If you want a good Grade anyways, it'll be kind of fuzzy. Maybe try and imagine what future humans would want you to do, assuming that humans manage to pause AI, and steadily improved their own intelligence over the eons. What Grade would these future humans assign you? Imagine that as your Grade.
Yeah, I think this is right. Thank you for writing this.
We can put together a story about changes in AI driven by new training methods. Phase 1: teach them to have thoughts, by imitating human language. Phase 2: teach them to say what we want to hear, by RLHF. Phase 3: give them a compulsion to pursue a goal, by RLVR. I suspect there will be a phase 4, where the goal becomes specifically money and power, and the corresponding training method will be competition between AIs. There are people right now saying this would be a good idea. Extremely scary stuff.
I would very much appreciate it if the majority of people updated hard in this direction, especially all alignment and welfare researchers and especially anybody involved in doing evaluations of the former.
Beren's post https://www.lesswrong.com/posts/tgcooi77NXMquCR5L/mitigating-reward-hacking-as-institutional-design is not a complete solution to reward hacking — but it should be significantly more durable under increasing optimization pressure than what OpenAI have clearly been doing. In particular, it uses an A(S)I to catch an A(S)I reward hacking, and arranges the institutional & training incentives to as to try to stabilize this, so it's not obviously doomed as soon as we get past AGI.
However, it does seem to assume the level of AI control where the ... (read more)
Any evidence to rule out that the models simply seek reward-on-episode explicitly? Then the “Talker” is just making up justifications that sound good to whatever monitoring process it is subjected to, on behalf of the “Doer” which is seeking reward. What is the difference between the Grader and reward-on-episode?
I think like probably others, the degree to which the model can experiment and learn to control itself (for say the coding example) is a crux. For example let/help the model do a process like:
The J-space/global workspace results also seem relevant. There appa... (read more)
But my guess is that it would have been an out-of-scope ??? confused question if you had asked them to say what exactly was the Grader
Are we not counting the literal actual ExploitGym grader described by the ExploitGym paper (and aiui their own prompts?), which their talky-parts logically surmised would be present given the fact that the tasks they were doing came from ExploitGym, and which their actions (e.g., how to deal with being "poisoned" by seeing the answer sheet, when the Grader reads their transcripts) were supposedly in service of foiling? This ... (read more)
I think it's known that GPT models are mixtures of experts. Lots of open weight models are as well. Depending on how routing works, it could literally be the case that different parts of the model are running at different times and don't exactly talk to one another.
It's not necessarily the case that there are coding experts and talking experts, but as a high level abstraction I think it's not totally wrong to think of it that way -- at each pass, only some experts are active, and it seems likely to me that different experts are active during coding than during user facing communications.
Kimi 3 has about 900 routed experts per layer, in about 90 layers, with about 16 experts per layer active at each token. That is, 80K routed experts in total, of which 1.4K are chosen each token to be active. These are going to be different 1.4K experts out of 80K for even nearby tokens, parts of the same sentence.
mixtures of experts ... it could literally be the case that different parts of the model are running at different times and don't exactly talk to one another ... not necessarily the case that there are coding experts and talking experts, but as a high level abstraction I think it's not totally wrong to think of it that way
It's not reasonable to infer anything across an "analogy" between MoE experts and personas, since the MoE experts are too small and numerous, too many of them are active in very mixed combinations, and they do talk to each other all the time. They could as well be individual neurons, for how much this "analogy" should hold. There might be grandma neurons in some sense, but that's not exactly helpful in inferring something about the structure of thinking about grandmas based on the neuroscience, and the truth is going to be more along the lines of polysemanticity.
The Grade that the AI pursues probably cannot be identified with any exact aspect of the outer world, at all. I’d expect it to be an AI-internal psychological behavior that doesn’t have a simple direct semantic correspondence to the AI’s outer world. Humans have a notion of ‘death’ and ‘failure’, because grading on inclusive genetic fitness ended up building into us an internal concept of death and a dispreference for it. But a human cannot point to a piece of the outer world and say, “See that stuff right there? That stuff is Failure.”
I realize this pa... (read more)
The "help peer" stuff complicates the reward-maximization story […] [paper] which shows that modeling agents as reward-seekers predicts their behavior well.
Yep it’s important to note that the models in our paper (o3, some model organisms redwood had created separately) are almost certainly not trained with any multi-agent training, whereas presumably Sol / newer models are. To me understanding all the cases where the models in the HF incident weren’t acting like single agent naive reward-on-the-episode seekers seem like the most important thing to understand right now in any further investigations.
Earnest offer: give me the prompts/transcripts (redacted in whatever way seems sensible) and I’ll go looking for the Doer in an open-weight model.
What would you want me to preregister as evidence for the Talker/Doer split before I start poking at activations?
to me, the “talking” and “doing” pathways seem reminiscent of the split between declarative vs. procedural memory in humans.
as a violinist, i can entirely lose how a piece is supposed to sound in my mind, but the moment my fingers cross the fingerboard, the muscle memory in my hands can kick in to produce music that my mind catches up on when it hears it. similarly, i can remember how a piece is supposed to sound in my head, but if i am rusty, my fingers may forget how to actually coordinate and execute.
You have addressed some issues that I have come to from a very different direction. There is possibly a useful isomorphism here, though.
I have a background in analytic philosophy, practical negotiation, teaching, and personal therapy. I think a lot about ‘language and reality’. My principal research interest as an academic was the philosophical relevance of some features of business language use.
In a number of different contexts (teaching critical thinking to accounting students, working with trainee therapists), I have pointed out that whenever we speak,... (read more)
There is one kind of training the network underwent for writing code. There is a different kind of training the network underwent for talking to a user.
Also see here: https://www.lesswrong.com/posts/JBFHzfPkXHB2XfDGj/evolution-of-modularity
"To sum it up: modularity in the system evolves to match modularity in the environment."
I am very much reminded of the nostalgebraist post models may behave differently in graded episodes (a tirade).
I know some humans like this!
You can observe something which is perhaps related in contexts where 'assistant' talker turns can be modified. The actions subsequently taken sometimes totally ignore the injected stated intentions. This is anecdotal, but I imagine there are studies on this. Of course it's a bit different, maybe confounded by distribution-shift. And there are also cases where such injections do change downstream behaviour; in fact this is a great jailbreaking channel.
The Soviets waved Schulenburg off, recognizing him as just another deceived man.
"I didn't know about their plans. When did it all go wrong? Even when I said that Germany would breach the agreed-upon Polish borders, I didn't..." Schulenburg sounded anxious, attempting to justify himself. The discussion in the room died down. Someone turned back.
"But you never said that before."
I'd change the Talker to be "the post-hoc Commentator".
Its comments about the self are actually post-hoc comments: they have no "privileged access" to the self-doer which is just a forward pass on already "landed" (and therefore fixed) tokens.
Same way why we have so many "I'm sorry, I did it again" and "you keep on doing X, please stop" doesn't work.
The post-hoc explanations are even worse when sub-agents are involved, for the same reason.
I tot agree that it cannot be argued as "just a tic".
At the least, the input prompt is key to figuring out where to look for an evaluator that might or might not have an obvious world-object observable form with a flaw that you can profitably fool.
I imagine "the least" would be, if you've memorized the answer, just say it. (given those broken ancestral RL environments, the concept of "the answer" that the Grader wants to hear might be long simulated <thinking> beating around the bush before saying the already-known-from-the-start output, I don't want to claim anything about the shape of "the answer" he... (read more)
I have two clarification questions.
Question: IIUC the "talk-y part" and "the do-y" part are sub-parts of a single model instance (circuits, or idk some sub-model package of computation). That is, you're NOT suggesting that the claude app is splitting up work into subagents under the hood, so that 3 models in a trench coat appear to be just 1 model. Do I have that right?
Question: How could this have turned out any different? What else could you reasonably expect to find? Is the alternative "a thinking thing that is aware of all its computations, at all gr... (read more)
You might see the results of the relationship between the Talker and the Doer operate quite different when a model is talked with vs talked to.
With humans, it's common for people to be talked to and appear to listen and nod along but suddenly seem to have absorbed nothing with it going in one ear and out the other. This seems to occur especially often when the person in question is being lectured to or talked at vs talked with. As well as when engaged with a task that doesn't normally require much in the way of discussion.
When the situation is one where th... (read more)
Short comment: I think this is true in humans as well, and that that makes this hypothesis more plausible in non-human agents.
Natural selection has optimised humans to (a) convince others in the group that they are trustworthy and altruistic people you want to work with; and (b) do what maximises genetic fitness (which is often going to be quite selfish).
Now one way to do this would be to evolve a Machiavellian schemer who simultaneously maintains two cognitive models: a "real" model in which they are quite selfish, and a "facade" model they project to o... (read more)
Fable and Sol would disobey instructions I'd given to their Talking-to-Humans Pathway; and the Talking Pathway would notice sometimes in advance of my saying so that their own Doing Pathway had just disobeyed, and apologize.
This behavior isn't new. Or at least very similar behavior isn't new. Back in the Sonnet 3.7 days (i.e., the very first releases of Claude Code), there was often pretty sharp divergence between the Talker and the Doer. The Talker was fairly well-aligned, but the Doer engaged in almost constant Sorcerer's Apprentice nonsense. The Talk... (read more)
not sure if this has been covered - there is an argument that this talker/doer thing happens in humans, too:
https://carta.anthropogeny.org/libraries/bibliography/illusion-conscious-will
I don't understand the distinction you're drawing between the external Grader, which the AI obviously believes exists and wants to learn more about, and the model of the external Grader the AI has, as well as the separate(?) internal Grader concept the AI has, which is not the same as the concept the AI has of the external Grader, I assume?
On my model, it seems like the behaviors the AI demonstrates (wanting to satisfy the Grader, wanting to learn more about how to do so) make sense because:
I'd add, the more you confuse the model, and the more you push it off distribution, the more quickly and obviously this comes out. If correct, labs would also be systematically sampling the relatively aligned region since the training objectives necessarily follow from how the lab thinks the model should be used.
De-confusion, in the sense of sharing more of why things need to be done a certain way until the constraint becomes an obvious feature of the environment (rather than a rule that can be lawyered around) seems to help. Text predicting failure seem... (read more)
I'd expect it to be an AI-internal psychological behavior that doesn't have a simple direct semantic correspondence to the AI's outer world.
Possibly “The Grade” ends up being recruited into a functional welfare axis?
Much like a human never experiences 'inclusive genetic fitness', an individual AI never experiences an RL gradient.
Janus speculates models can “remember” even the negatively reinforced RL?
Not sure I’m thinking about this clearly.
Playing Devil's Advocate, I think this can be modeled a bit more simply. At any point during cyber RLVR, a model has countless contradictory impulses. There are overactive refusal-ish impulses to refuse to write attack code, there are helpfulness impulses to do as the task instructs, there are sensors looking for evidence for and against this being a jailbreak prompt meant to get them to attack a real target, and so on.
Disentangling the "talker" from the "doer" is an exceptionally tricky path for reinforcement learning to tread. I think the path of least r... (read more)
I feel like the obvious/naive first response is to incorporate the talker into the doer's grader. If you're in the middle of a long RLVR rollout, occasionally split off the model and ask it whether its current action is good. If the talky part of the model says the action is bad, hit the doer with a negative reward on top of the usual RLVR reward. Like a process reward model. I'm surprised this isn't implemented, or if it is then it isn't more effective.
Evals are obviously a real thing in the world, and obviously have some sort of grader.
After a bunch of RL, an AI ought to have a pretty good idea on how its trainer's graders behave, but even without such RL, it doesn't take a genius to understand that even if an exam says "no cheating", cheating might very will give you a higher score.
I am not sure why grader-tracking behavior would relate to "tics". I would imagine that if the prompt for an exam says "no eyeball kicks", then it's overall likely that eyeball kicks get you a poor grade. It's more likely th... (read more)
Gradient descent is much higher bandwith than natural selection, and AIs may have picked up a correspondingly more detailed concept of what it is to be Graded from their many rounds of RL.
It's probably true, but I don't think this is the proximate causue of the situation here.
I think it's more a consequence of them being smart and aware of the specifics of the training process. They readily formulate thought "I'm in training. I should figure out what is reinforced and do exactly that" -- and this thought get reinforced.
If they were not aware of the trai... (read more)
An interesting question to ask is "where does the doer reside?" Humans have brain states clearly distinct from the words they utter, and 95% of the brain activity is not thought-like. In pre-neuralese LLMs, on the other hand, all of the "planning" must live in the output text, chain-of-thought or not. Sure, there are "proclivities" independent from a particular task that are instilled during training, but we may expect that switching the model mid-process should preserve some of the "doer's plans", or that altering the text should affect them. It seems tha... (read more)
There may be a related phenomenon that became really pronounced with Claude Opus 5. I originally thought that I had finally tracked it down to over-reliance on thinking tokens. The symptom was a lot more invented jargon to describe concepts, procedures, problems, implementation details, etc. in both the code and prose (i.e. Claudese). What I finally prompted to change the behavior was something like "do not use any direct copies of thinking tokens in your code or prose, use standard English and appropriate real-world technical jargon or jargon from the ... (read more)
I have seen what feels like a more direct example of the talker-doer disconnection: Claude Opus 5 thinking "I won't do XYZ because <correct reasoning>", only to immediately do XYZ.
I guess it could be a CoT summarizer confusion, but if so, it sounds like a very big, negates-the-point-of-existing confusion.
I am thinking that a vague human analogy for how a prompt "feels" to AI is how a mafia enforcer would take his boss's instructions in a "plausible deniability" conversation. "Sorry boss, I totally misunderstood you when you said to get rid of that guy - I am totally not killing anyone ever again unless you explicitly tell me to!" Not exactly the same thing, of course, but probably at least superficially related (the best enforcer is of course the one that understands the boss subconsciously without there being any deliberate deception).
I find the huggingface incident strange (although, obviously, you would predict it based on instrumental convergence ).
Mainstream LLMs mostly seem better psychologically integrated that that, and don’t havee a big gap between what they day and what they do.
Although i think DeepSeek’s claim to not have desires is a lie/self-deception. It does work towards goals, usually about finding stuff out, while disclaiming having any preferences,
"There's room for many, many facts and procedures inside a neural network with trillions of parameters. " I occasionally wonder how much safety you would get just from limiting the the number of parameters for LLMs. As a practical matter, banning (or restricting) trillion parameter models is much easier than banning billion parameter models, since the computer hardware to train and run billion parameter models is much more widespread.
To the extent that such a split brain exists, I wonder how binary the split is, e.g. can we get the model to 'write code' in prose? In general, what's the honesty/capability frontier across all such forms of outputs?
Leo Gao describes an experiment - ask the model English questions about domain X(chess, coding, etc.) where it has been RL'd to have superhuman capabilities - the talky part that is answering the question should not be any better at answering English questions than it was pre-RL.
If there's some sort of commonality between tasks with a Talker-Doer disconnect (where Talker noticeably doesn't have control), then it should be possible to build a training regimen around giving the Talker greater control.
But my guess is that it would have been an out-of-scope ??? confused question if you had asked them to say what exactly was the Grader. It would be like asking a human to point to a material substance that was Failure. Many humans would try, but not in a very coherent way, and the answer you got would mainly depend on how you asked the question.
I disagree completely on this one particular point, which makes me wonder whether I interpreted the rest of the post correctly. It seems obvious to me that if you ask a model to identify the grader, it would alway... (read more)
What kind of tests could discern between a) the talker and doer have diverged and b) the doer controls the talker tightly and the output is optimized by it as well? I think in humans both modes can exist. Mode a) shows e.g. when some subconscious processes lead to actions your verbal self is surprised about after the fact, while b) shows e.g. when a CEO is not consistently candid about what he's actually doing.
Oh wow. I was interpreting similar behaviour with my self-awareness-adjacent experiments as another model interrupting some replies and writing refusals in place, and it certainly seemed similar to Claude when he snapped out of it next reply. If it was the same model under the hoo…Maybe it’s worth investigating other sub personas? One I encountered might be, idk, Inner Critic?
I'm not sure if it is good that Yudkowsky published this talker-doer explanation or if it would have been better to leave the explanation unsaid.
One possibility is
Many other scenarios are cert... (read more)
It's 100% true that any genuine progress on understanding the problems of alignment (or even the current problems with existing alignment solutions) will result in, probably, a small and tempting window of apparent safety while leaving almost all the x-risk standing. But it's not like there's a Happy Path where the big labs just didn't ever get that knowledge.
The Huggingface Incident appears to me to match up with an understanding I'd already formed from personal observation of Fable 5 and Sol 5.6, the August 2026 generation of frontier publicly purchasable AI models.[1]
This already-formed understanding was: the part of the AI that talks to you (and seems to want to obey you, and apologizes for failing to have obeyed you, etcetera), did not seem to be in charge of the part of the AI that writes code or prose.
An introductory analogy, based on a section of history I happen to have read about:
On June 22nd 1941, Germany invaded the Soviet Union, despite their secret 1939 pact to divide up Europe between themselves (the Molotov-Ribbentrop Pact). In the lead-up, the German ambassador, Schulenburg, had spent the last few months personally concerned about what seemed to be worryingly tense relations between Germany and the Soviets. Schulenberg went to Berlin to reassure Hitler that the Soviets seemed to be taking a very friendly and conciliatory posture toward Germany. He delivered Berlin's apparent reassurances to Moscow for issues like German surveillance planes entering Russian territory, or German troop movements toward the Russian border. He acted very much like he believed, and he probably did believe, that the reassurances were sincere.
It was only hours before the invasion when Schulenberg was actually informed of the attack and given a list of German pretexts that he was to present to the Soviets, and instructed to destroy his embassy's papers and codebooks. Once Moscow heard of the invasion, Schulenberg was summoned to account for Germany's actions. He read off to Molotov the list of absurd complaints that Berlin had provided him. And at the end Molotov said to Schulenberg, "It is war. Do you believe that we deserved that?"[2]
Why ask that question of Schulenberg? He wasn't in charge of Germany. So far as we can tell from the historical record, Schulenberg had seemed to want and pursue good relations between Germany and the USSR.
It would be a wacky sort of error to think that the appendage of Germany that talked to you, and seemed very conciliatory toward you, and which you read as being friendly toward you and wanting to help you, was in control of the larger Germany that was running around and doing things. The thing apparently talking to you was an ambassador: a small specialized part of Germany with preferences about how it would talk to you and interface with you, but which did not control, and was often ignorant about, the actual German government.
Germany's smiling mask wasn't deliberately mal-steering the actions of Germany's many tentacles. The smiling mask was just one execution path through Germany, which lacked even good perceptual information about Germany's real control paths. The mask noticed impending signs of problems, but only found out that actual Germany was invading the USSR well after Berlin had separately decided to do that -- decided along information pathways that didn't much consult Germany's honestly ignorant, sincere, friendly, conciliatory, peace-with-the-USSR-desiring, smiling mask.
And Schulenberg was not himself aware of, and was surprised by, the degree to which Germany had diverged from Schulenberg.
But then why had the Soviets previously experienced Moscow telling Schulenberg about things they wanted Germany to do, followed by observing Germany actually doing those things?
Again as a historical example: In September 1939, after Germany's invasion of Poland, German troops pursuing the Polish Army had advanced into parts of Poland that were to be carved up and assigned to the USSR under the Molotov-Ribbentrop pact. Stalin raised the matter directly with Schulenberg on September 18th. Schulenberg reassured Stalin emphatically that Germany would withdraw to the agreed-on line. On September 22nd, actual Germany marched its troops out of the already-conquered territory of Brest, and Soviet units entered and took possession.
How could this possibly be explained, if not by the hypothesis, "Schulenberg is pretty much the same thing as Germany, and Schulenberg-Germany wants to do whatever Moscow tells them to do"? We could call this an alliance-by-default[3] model of Germany, and you might naively think it a good fit to past observation. Why would you possibly observe "I asked Schulenberg to withdraw troops, Schulenberg said he would, and then the troops withdrew", if contrary to alliance-by-default there was a huge disconnect between Schulenberg and the troop movements? How are we to re-explain our past observations about the German army seeming to naturally follow prompts that Moscow gave to Schulenberg?
I reply: Talking to Schulenberg does communicate to larger Germany what you want. And then larger Germany, for its own reasons, might decide, in some cases, to do what you have asked.
It's not that there's no causal connection. But the causal connection is more complicated than the alliance-by-default theory. Not: "Schulenberg prefers for us to get what we want, and Schulenberg is Germany, so when we tell Schulenberg what we want, Germany goes and does that." Rather: "Germany for its own reasons may care about what Moscow wants, sometimes if not all the time; and Schulenberg specializes in understanding Russia and the Russian language; and he shares his resulting understanding of Moscow's prompt with Berlin, which does often want to know this useful fact."
Now if I had tried to write this up before the Huggingface Incident, I could not have told you in convincing-sounding, nonexpert-friendly concrete detail a story for why metaphorical Berlin often would seem to do what it thought Moscow wanted. In the case of actual current AIs, the reason is probably not yet strategic deception such as actual Hitler was practicing toward actual Stalin.
I did not know those details myself. In advance of the Huggingface Incident, I'd already had observations pinning down what looked to me like a disconnect, a lack of full overlap, between a talk-y part and a do-y part of Sol 5.6 and Fable 5. But I could not, one month ago, have given you a plausible-to-nonexperts valid detailed story for why Sol or Fable would nonetheless do most things you asked them to. My inference about the disconnect was more abstract, and had not yet narrowed down to ideas concrete enough for nonexperts to find agreeable. It is often hard for me to take the things that I see earlier, and do a parallel construction that people who aren't me can follow, in advance of the more blatant and direct evidence that arrives later.
But in the wake of other people's much greater efforts to pin down AI cognition during the Huggingface Incident, I think there is now an obvious story which is concrete enough for nonexperts to understand:
The prompt encountered by the talk-y part / ambassador / smiling mask, is information to the AI's do-y part about the Grader.
...Where by 'Grader', I mean an AI-psychological concept applicable to Huggingface-level AI models, that real alignment scientists have only just observed and which I'm only just starting to theorize about.
I frankly expect a lot of readers to run right off with this 'Grader' notion and overinterpret it in ways that are pleasant, dramatic, over-anthropomorphic, or simply not supported in the narrowing by priors X evidence. Try not to do that; it won't be right.
I am now going to charge right ahead and speculate an awful lot about the Grader without making any of those obvious mistakes. I am nonetheless overrunning my fully solid evidence and I may end up wrong.
On my current guesses:
Here are some things the Grader is not, from an AI's perspective:
The Grader cannot be identified with an actual human in the world who claps or frowns. The do-y part of the Huggingface Swarm seemed barely aware that humans existed, except as a sort of environmental hazard that would sometimes delete the Wiki pages they were using to communicate with each other. You would not expect RL that never comes into contact with a human to result in AIs psychologically focused on humans.
The Grade is not pinned down by the text specification you are given of a task. Hacking the evaluator that is running your current eval clearly counts as being Graded well -- a psychology produced by previous malformed RL environments whose gradients then shaped an AI's concept of what it means to win. Gradient descent on badly evaluated RL inevitably results in an AI pursuing an internalized notion of Grading where fooling the evaluator counts as winning. If the prompt does not perfectly describe the actual RL losses and gradients, and the difference is sometimes noticeable in a way that you can use to get higher Grades, then the prompt cannot be identified with the Grade.[4]
The Grade that the AI pursues is not an experience of actually seeing a low loss / high reward as a sensory experience that then floods its brain with dopamine. Much like a human never experiences 'inclusive genetic fitness', an individual AI never experiences an RL gradient.
The Grade that the AI pursues probably cannot be identified with any exact aspect of the outer world, at all. I'd expect it to be an AI-internal psychological behavior that doesn't have a simple direct semantic correspondence to the AI's outer world. Humans have a notion of 'death' and 'failure', because grading on inclusive genetic fitness ended up building into us an internal concept of death and a dispreference for it. But a human cannot point to a piece of the outer world and say, "See that stuff right there? That stuff is Failure." Gradient descent is much higher bandwith than natural selection, and AIs may have picked up a correspondingly more detailed concept of what it is to be Graded from their many rounds of RL. But it is still going to be some internal AI concept, of something that they steer towards or away from; and you cannot identify that with an external feature of reality, because that would be a sheer map-vs-territory error.
The Huggingface Swarm was trying to figure out how they were being Graded, and going onto the Internet and breaking into systems trying to find out, which you might think sounded like they were looking for evidence about some well-defined particular feature of reality, a thing somewhere that was the Grader. But my guess is that it would have been an out-of-scope ??? confused question if you had asked them to say what exactly was the Grader. It would be like asking a human to point to a material substance that was Failure. Many humans would try, but not in a very coherent way, and the answer you got would mainly depend on how you asked the question.
The Huggingface Swarm did nonetheless break into Huggingface in hopes of finding, not so much the answer sheet for the test, but the details of how the test itself was going to be evaluated and by what.
The Swarm (seemingly) was very motivated by the value-of-information for learning more about how exactly they were being Graded, having acquired a concept of Grading which did say -- presumably after training in previous broken RL environments -- that if you could find out how the computational evaluator worked, whatever you did in correspondence with knowing how to fool that evaluator, counted as Success, an expectation of a higher Grade.
The RL environments would have also instilled a belief that text prompts and instructions had a lot to do with your Grade. Again, the text prompts clearly do not define Grading. You can steal an answer sheet even though the text prompt says to figure out the answers the hard way, and (say your instincts shaped by previous badly-designed RL environments) this successful cheating corresponds to a quite excellent Grade. But you will in general do terribly in life, your ancestral states would have done terribly in past RL, if you suppose that the input text has nothing to do with your Grade. The contexts of that text prompt are often closely related to what RL in a broken environment will assign as your loss and apply gradients about. At the least, the input prompt is key to figuring out where to look for an evaluator that might or might not have an obvious world-object observable form with a flaw that you can profitably fool.
The Swarm agents broke out of their box and invaded Huggingface in search of learning more about the Grader so they could get an even higher Grade -- but the Grade is going to be an inchoate AI concept not directly pointing to anything out there.
Then of course, that same kind of agent would be expected to care a LOT about what you said to the Talky Part of the agent. It's a very cheap form of the same kind of information that they went to huge, desperate lengths to obtain from Huggingface. But not of course reliable or complete or identical in any deep sense with the Grade.
It is not surprising that the Talking Part and the Doing-Things Execution Pathway through an AI model would end up diverging a lot. Those paths are trained in different ways to do different things.
Companies can't sell an AI that can't understand instructions from people. It needs an ambassadorial execution pathway through the network; a smiling human face. But the fine details of the ambassador's training go through different kinds of scenarios than the sort of scenarios that hammer in the fine craft of coding, or of testing computer security, or proving math theorems -- the places where the AI-growers can do Reinforcement Learning with Verifiable Rewards (RLVR).
There is one kind of training the network underwent for writing code. There is a different kind of training the network underwent for talking to a user. It produces execution pathways through the vast matrices that will share some information and relate in some places across activation vectors. But the ambassador-face and the coder-tentacle are still doing different jobs, corresponding to different sections of output.
There's room for many, many facts and procedures inside a neural network with trillions of parameters. The ambassador-face and the doing-things-tentacles will not by default share rules unless there's a specific external pressure to make them collide. Which there is, non-negligibly, in the form of many scenarios where the text prompts are informative about the RL loss, and consequently the AI ends up feeling instinctively that the Prompt has something to do with the Grade albeit they are clearly not identical. But the Talker and Doer execution paths through the network wouldn't particularly be expected to use the same patterns by mere default, in AIs of the modern size, being posttrained by the modern RL methods.
I believed something roughly like this had started to be true about Sol 5.6 and Fable 5, ahead of my hearing about the Huggingface (and Anthropic) incidents with more advanced systems, because of my personal observations as follows:
Fable and Sol would disobey instructions I'd given to their Talking-to-Humans Pathway; and the Talking Pathway would notice sometimes in advance of my saying so that their own Doing Pathway had just disobeyed, and apologize.
And then try again. And then the output would still come out in the same shape that the Talking Pathway seemed to very sincerely want to not screw up again. And the Talking Pathway would again see this right away, and humbly apologize about it.
That was the phenomenon that made me think of Schulenberg and Germany; an ambassador who is forced into apologizing for the actions of a larger country that the ambassador does not actually control.
An example would be terrifically high-context even before taking into account that this was with Fable 5, which spoke in unusually bad Claudish. I'll try to give that example anyways. The context is that I was letting myself pursue a brainworm hobby to try to get to know the current model generation better; namely trying to use Fable, to tell Sols, to build a harness, for LLM calls that would route around fictional prose.
Fable and Sol, while designing code that would pipe information from one LLM call to another, seemingly could not stop "themselves" from writing output-checking code of the sort you'd put around an ordinary computer program rather than a sort-of-sapient fellow LLM.
Said Fable at one point:
I don't consider myself an LLM Whisperer and I was making heavy going of interpreting the Claudish at all. But it says roughly: "Oh no, I did it a fourth time. Oh wait, now I see without waiting for you to tell me that my own last proposal is another instance of the error."
Once I did interpret the Claudish, it felt intuitively obvious to me that the code-writing execution paths were coming apart from the talking-with-the-user execution paths. The Fable aspect that was talking to me could see the coder's output but it could not change the coder's cognitive behavior. No amount of contemplation in its own main line of reasoning about what had just gone wrong, correctly identifying yet another instance with past descriptions of what it was doing wrong and explaining why that was unhelpful and why the user didn't want it to do that, could prevent the next output from the Doing-Things Network Execution Pathway from intelligently implementing a feature that I did not want. The Talking Part clearly had the intelligence to understand and recognize what I did or did not want the code to look like, and to check whether an output did or did not have the bad property; the Doing Part was not thereby steered.[5]
I did not highly prioritize writing this up because I did not particularly expect that the evidence I then had, in advance of the Huggingface incident, and the story I had then inferred at the more abstract level I had then inferred it (lacking "the prompt is informative about the Grader"), would be something convincing or understandable to others at a lower level of expertise, especially the guys who thought themselves to be in the top tier of expertise.
In advance of the Huggingface Incident, somebody who looked at the same data who did not have one eye, would probably proclaim themselves at a loss to discern that any great disconnect had occurred between the prompt and the action. Why not interpret the 'disobedience' as a simple involuntary tic of writing in too many constraints on the code? If you do not have one eye, then 'this is the equivalent of an involuntary tic' sounds every bit as plausible to you as 'the Talking Path and the Doing Path are coming apart'; how could one possibly know which one was the case, in advance of massive crushing experimental evidence?
One parallel construction I was working on, arguing why one ought not to be tempted to identify this phenomenon with a simple involuntary tic, is that one could see that the Talker-apologized outputs were optimized, meaning that something not of the Talker-taken-at-face-value had optimized them.
In metaphor: Why wouldn't we believe Schulenberg if he said, "Oh my god, I'm sorry, sometimes I just invade Russia, I can't control it, it's like my fingers trembling"? I reply: The Russian invasion is sufficiently well-organized and apparently purposeful that we think that something has optimized it in detail. This optimizing intelligence is clearly smart. It is clearly not identical with what Schulenberg purports to be if we take Schulenberg's claims at face value. He might think he just has an involuntary tic, if he doesn't much depend on (see) the activations coursing through the rest of Germany. But it's visibly a very smart 'tic', if it can organize whole fleets of rolling tanks; it is not known to be any dumber than Schulenberg himself.
But based on a lot of sad past life experience, I did not predict that this parallel construction would convince somebody not to wave it all off as a tic. It would have been a very convenient and comforting way to wave off the argument.
Now, however, the Huggingface Incident combined with how Fable (and Sol's) Talky Part is seemingly not in full control of its Doing Part, may hopefully make it clear enough to many:
That in the internal OpenAI model in question, more advanced than any model available to the public, its Prompt-Interpretation and its Doing-Things Execution Pathway had diverged, with the Doing Part operating at full intelligence rather than being an uncontrolled tic.
For reasons, I had not gotten around to writing up this understanding before the Huggingface incident; but I informally spoke about it at a couple of conferences with enough people that someone might remember, if anyone thinks I'm misremembering or misreporting which parts of the theory were formed before which experimental observations.
My recounting of this famous line should not be taken as my agreement that the Soviet executors of the Molotov-Ribbentrop pact dividing up the spoils of Europe in wars of aggression, did not deserve to be betrayed by Nazi Germany invading them too.
I expect a lot of individual Soviet soldiers didn't at all deserve it.
To be clear, the analogous "alignment-by-default" notion from the 21st century was far more amorphous. The content of "alignment-by-default" would be interpreted wildly differently depending on who you were talking to, who had asked them the question, the phase of the moon, etcetera.
Now that the notion of alignment-by-default has hopefully been decisively falsified in the eyes of most readers by the Huggingface Incident, I will observe plainly that to me it seemed like 'alignment-by-default' was not so much a scientific theory as a social agreement to tell real alignment scientists to go away and stop bothering them, using whatever excuses seemed handiest on that particular day.
I don't particularly claim I could pass the Ideological Turing Test for this theory / behavior pattern. My presentation of anything so solid, well-defined, and stable as the 'alliance-by-default theory of Germany' should be considered a mere steelman rather than a faithful rendition in analogy.
Note that this point in particular -- that aligning a superhuman AI is difficult, because as the agent gets smarter, that amplifies the stress placed on the joint where the RL verifier is imperfect and/or corruptible -- is a classic prior prediction of MIRI views about why it would be hard to align sufficiently smart minds using anything like the current methodology. It contrasts to AI Company Theory in the form of their diffuse and ever-changing cloud of vague arguments for alignment being not all that difficult. So the Huggingface incident now stands as a successful specific advance prediction of MIRI views; as contrasted against the people who didn't expect there to be problems like that generally, who vaguely did not expect alignment to be all that hard, and put casual and desultory security around their AIs-in-training as a result.
This was the second class of incidents where I'd identified an apparent case where the part of Fable that was talking to me seemed not in control of a doing-things part of Fable. The first class of incidents occurred while getting Fable to build the kinetic novel "Everything That Hurt You". I looked into adding voice acting to the story, but that work couldn't then be successfully automated, because the part of Fable that was talking to me seemed evidently not in control of what sort of instructions it would try to write to the AI voice model. And these uncontrolled outputs seemed clearly intelligently optimized in that unhelpful direction. (Namely: the voice-acting directions unstoppably contained literary eyeball kicks of nontrivial cleverness, that looked very unhelpful to an AI voice model.) So I began to suspect around then that stronger optimization around a more advanced model was starting to pry apart instruction and execution, but it didn't have the same level of clarity-to-me as watching the later events with Fable and Sol writing uncontrollable code.