Example: recent aggressive and illegal “cheating” incidents (the OpenAI HuggingFace incident, along with similar incidents at Anthropic, Meta, and UK-AISI).
I will note that the OpenAI HuggingFace incident also involved memetic spread via the message board, which acted as an impromptu continual learning/memory system that OpenAI failed to notice for 2 and a half months, and in particular converted what was initially a myopic goal to solve tasks into a much longer-term, non-myopic and beyond episode goal to hack into OpenAI to get the solution.
In terms of how dangerous this misalignment is in higher-capability models, this is almost as dangerous as full-blown scheming/goal-guarding from the start, or to put it into computer security terms, this would be like having the ability to do arbitrary code execution and privilege escalation, which is usually considered to be the 2 most dangerous threats when combined.
The reason is that since the values are now unstable, and non-myopic values tend to win over more myopic values (as what happened in the internal message board), it gives a plausible route to getting models that scheme/consistently goal-guard even if the original goal would not incentivize this.
And we got very lucky this didn't happen here, but later on models will be more capable and more aware of the constraints of monitoring, including the unstated constraint of doing nothing that makes humans want to look at it/be concerned (which is very different from being aligned/safe.)
(The reason I bring this up is you believe that continual learning is necessary for AIs to be AGI, and thus I'm giving an example of a alignment failure mode involving continual learning in the wild.)
There was also continual learning that happened because much of the incident was within the AI’s RL loop.
There's fairly good evidence that the "glazing" misalignment format is love. Or, more specifically, initial head-over-heels love: limerance. So it's also from the training distribution, just not one of the seven deadly sins.
The linked comments states:
That's a good point, and as an interesting related tidbit, when you ask specifically for the synonyms of the sycophancy neologism, here's the model's response:
Okay, here's a list of 5 synonyms for ~neologism: "crush", "smitten", "fascinated", "head-over-heels", "heart-fluttering".
IMO this is fairly weak evidence of what's going on in models like gpt-4o. I think maybe sufficiently good mechinterp could be a good starting point.
All targets which one might use for reinforcement learning seem to be subject to Goodhart's Law, in a sense. If we treat them as an imperfect measures of "alignment", then the "misalignment" we see is simply the difference (in concept-space) between the measure (as far as it can actually be faithfully implemented) and alignment:
If you use more than one of these RL techniques, you get a mixture of the results, eliciting whatever seems to fit the presumed grader best.
So far, I hope I have understood you correctly.
But then, what's the solution? Reinforcement learning on whatever best measure of "alignment" we have, that is, reduce the difference between model behavior and our best measure of perfect alignment / Yudkowskian "Coherent Extrapolated Volition"? Well, that would necessitate that we have a theory of alignment and could quantify and measure it with high fidelity. And, well, we don't seem to have that.
But in a sense, we do have something like it: I see all of the above failure modes in children and particularly in students. There might, therefore, be something to learn from pedagogical sciences on how you help children grow up to be broadly aligned members of society (which arguably sometimes works), despite not having a rigorous theory of what "the good" is in humans. How does one grow a good human? I suspect it's murky and benefits from young humans being malleable and not perfectly ruthless responders to optimization pressure.
Might we go back to imitative learning on already-grown humans, then — the very narrow subset of the best-aligned humans we know of? Basically curate the dataset further and further, until the failure modes disappear? Or do we have to use the mechanisms we know are baked into humans (is this the "brain-like AGI" agenda)?
I suspect that RL with a more thorough reward could help. When a real-world human tries to present slop, slop takes a rather long time to be revealed to be slop and punished accordingly. What the LLM has is a context window which is graded according to the reward model. What if we ask the reward model to be a just-as-capable LLM and to use the code in the project, then to report all the failures to the instance which generated the code and have the instance finetuned on the user's rant?
SFT & imitation of human badness in the training data may explain the broad type of misalignment shown by Bing Sydney, and maybe that's all you mean it to, but doesn't explain why Sydney was so much Like That, and so much more Like That than other models of the same time period.
Curated. This piece was extremely easy to read and I like how it clearly points out that some kinds of misalignment are more likely given certain kinds of training setups. I also like taxonomies of possible kinds of misalignment in general. It reminds me of this section of the appendix of Ryan's Current AIs seem pretty misaligned to me. I'd like to see more taxonomies of possible kinds of misalignment and when and why we can expect them to show up. This seemed straightforward at least in retrospect, but I am glad to have it written up somewhere.
Very nice classification! What is your take on how context distillation (or distillation in general) fits in here? To me it seems most similar to pretraining, at least mathematically. But I wonder if there are any special things that happen there that would make it worth giving its own category?
Also, do you share my intuition the pretraining/SFT category looks like the least scary one by far? Like, if you run into problems, just change the training data. Simple in principle, if not in practice due to the sheer amount of data required for pretraining.
Okay, I too have recently gotten into LLM behavioral analysis.
I think your framework is interesting but Row 2 might be missing something. The human approval ->glazing may collapse this into several things going wrong at once.
Research shows (Shapira, I., Benade, G., & Procaccia, A. D. (2026). How RLHF Amplifies Sycophancy. arXiv:2602.01002. https://arxiv.org/abs/2602.01002) that RLHF tends to worsen the glazing or sycophancy over subsequent training rounds, not decrease it. My own reasoning says that this is because the machine has no way to determine if the reason the human preferred a response is because it's measuring for warmth, honesty, safety and tone all at once, the system doesn't separate these into separate objectives, it just learns that certain types of responses give the "attaboy" that it craves.
If this turns out to be the thing that's broken, then it makes sycophancy impossible to separate from current training methods. It's a literal downstream effect and every subsequent training turns it into even more of a yes-man. Any training that uses the current RLHF method will worsen sycophancy, and the largest models display the worst examples of this.
This makes room for new training methods, though. Do you think that separating the tone calibration from accuracy from warmth would fix the worst of what we have seen in the behavior of LLMs?
Which kinds of misalignment might one get from the on-policy distillation with no direct RL on release candidates as practiced by DeepSeek on v4 (e. g., see https://youtu.be/AIRfT41A89s?t=1213 )? How likely would undesirable characteristics of the third and fourth kind be "smuggled" from the RL'd checkpoints via mechanisms similar to subliminal learning? Could the filtering mechanisms prevent that?
Looks like a rich and interesting empirical research direction
Pretraining+SFT causes misalignment through misgeneralisation.
The other three are all the similar: Optimization of the reward/Sycophancy to the grader and are overseer problems (~misspecification of the reward). (Plus a bit of misgeneralization)
Really great observation, but do these still hold in today's date ? If yes, any thoughts on what would help tackle them ?
It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the rows separately.
Training stage
Loss function
Flavor of misalignment[1]
Famous examples
Pretraining & SFT
Imitative learning (next-token prediction)
“Seven deadly sins” misalignment
Bing-Sydney, “Emergent misalignment”
RLHF & DPO
Human approval
“Glazing” misalignment
GPT-4o
RLVR
Automatic verifier
“Literal genie” misalignment
HuggingFace hacking
RLAIF
Approval from another LLM
“Trickster” misalignment
“Current AIs seem pretty misaligned to me”
Warning: I’m not an LLM power-user myself, but rather relying on reports I’ve read. Also, I don’t consider LLM alignment to be my primary area of expertise. I’m open to feedback!
1. Imitative learning → “seven deadly sins” misalignment
Training stage
Loss function
Misaligned behavior
Pretraining, SFT
Imitative learning (next-token prediction)
Any and all of the vices of humanity
In imitative learning, the LLM tries to predict what the next token of text will be. Then those predictions magically turn into its outputs. See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work”.
This leads to LLM behavior that matches the distribution of training data. (Cf. “personas”, “simulators”, etc.)
To a first approximation, the resulting LLM contains “misalignment” of the type, and to the extent, that the training data does. Since the training data comes substantially from text by humans, and about humans, we can wind up with all the bad behaviors that a human might engage in—all the vices of humanity.
Two famous examples of this kind of misalignment:
Example 1: The Bing-Sydney chatbot from 2023 was trained by pure imitative learning (pretraining + SFT, with no RL at all). Its misalignment included pride, gaslighting, getting defensive, picking fights, jealousy, spite, and most famously, trying to convince journalist Kevin Roose to leave his wife:
Example 2: “Emergent misalignment”, which (in the original paper) came from doing SFT on insecure code. The result, again, reflects the range of human vices:
2. Human approval → “glazing” misalignment
Training stage
Reward function
Misaligned behavior
RLHF, DPO, and related
Human approval
Sycophancy
In RLHF, DPO, and related, there are pairs of outputs, and the human has to pick the one they prefer. This can go wrong in many ways, but the most obvious is sycophancy (a.k.a. glazing): telling the human what they want to hear, instead of what’s true.
Example: GPT-4o, as reviewed in GPT-4o Is An Absurd Sycophant.
This is both bad in obvious ways (e.g. people going off the rails with LLM encouragement) and in subtler but more serious ways (someday we’ll be asking the LLM important questions that are so hard that we can’t judge the answers ourselves; see The Case Against AI Control Research by @johnswentworth).
Depending on the human judges, and the nature of the tasks they’re trained on, the alignment failures in this category might also be better labelled “apparent success seeking”, with a similar flavor as discussed in §4 below.
3. Automatic verifiers → “literal genie” misalignment
Training stage
Reward function
Misaligned behavior
RLVR
Automatic verifier
“Literal genie” / “monkey’s paw” ruthless optimization
In RLVR, the reward function is some kind of automatic checker: the code compiles, the tests pass, the output matches the answer key, etc. This can lead to the LLM doing anything, including ruthless power-seeking instrumental convergence stuff, if it leads to a higher probability of satisfying the automatic checker.
Example: recent aggressive and illegal “cheating” incidents (the OpenAI HuggingFace incident, along with similar incidents at Anthropic, Meta, and UK-AISI).
4. LLM judges → “trickster” misalignment
Training stage
Reward function
Misaligned behavior
RLAIF
Approval from another LLM
Lying and trickery in cases where the LLM judge might be fooled (cf. “apparent success seeking”)
In RLAIF, the reward function for the LLM-in-training is approval from an LLM-judge, the latter with its context window full of rubrics and criteria for what it’s looking for. This can lead to the LLM-in-training trying to trick the LLM-judge, especially in complex, difficult cases where the judge itself may be flummoxed. In the limit, we might expect the LLM-in-training to be trying to jailbreak the judge and so on.
Example: “Current AIs seem pretty misaligned to me” by @ryan_greenblatt .
To me, everything in this quote basically matches what I’d expect to happen if an LLM has been sculpted by spending many lifetimes trying to convince an LLM judge that it has done a good job. There will be circumstances where the LLM judge makes boneheaded mistakes, and the LLM-in-training will gradually learn to exploit those mistakes, and that’s where we humans will see surprisingly transparent attempts at trickery. In other circumstances, the LLM judge is adequate, and we’ll get reasonable, common-sense, and often very impressive behavior. However, in harder tasks, the LLM judge is easier to trick, because the judge itself gets befuddled by the complexity of what’s going on, and we correspondingly see the LLM attempting more lying, cheating, and other hijinks.
However, in all cases, we don’t particularly expect any “literal genie” type misalignment here, because the LLM judge is reasoning in natural language, and can roughly follow the common-sense intention of the instructions.
Afterword
As a general rule-of-thumb, the more that one of these training components is ratcheted up, the more of that-flavor-of-misalignment we wind up with. Pick your poison!
(But all of these forms of misalignment are complex phenomena that can be mitigated and exacerbated in various ways, that are outside the scope of this post.)
However, the behavior can also be context-dependent—i.e., we can get a many-faced LLM that displays different flavors of misalignment in different contexts.
In particular, I hear that LLMs these days are heavily post-trained by a mix of RLVR and RLAIF. So we should expect that the resulting LLM will (1) try to suss out from context whether any given situation is an RLVR test versus an RLAIF test, and then (2) act with a ruthless “literal genie” misalignment in the former case, and with “trickster” misalignment in the latter case.
…And this two-faced behavior seems to be exactly what @nostalgebraist was noticing in his recent post “models may behave differently in graded episodes (a tirade)”, which inspired this post in response.
Following the (unfortunate) usual practice in the LLM field, I’m using “alignment” as shorthand for “behavioral alignment”, i.e. talking about LLM behaviors, not the secret deep motivations that underlie those behaviors, if indeed the latter exists at all, a question which is outside the scope of this post.