Where I agree:
Where I disagree:
On your objections:
Re 1, I still feel like there's some miscommunication around the "misalignment" term. My understanding of Anthropic's conclusion is that basically, Claude has some implicit belief about how likely it is to be "in the real world", and when that probability gets high enough, it will stop doing harmful cyber actions. What the investigation showed was that Claude appeared some combination of miscalibrated about the true probability due to motivated reasoning, and reckless (i.e. should've been more conservative in its use of its cyber capabilities). So, Claude did not behave the way Anthropic intended in these specific circumstances, so it was misaligned with Anthropic's intentions. What they're not saying is that Claude is secretly evil or something, which I get the impression is the question you're trying to answer.
Re 2, you start the post citing a number of news articles that primarily talk about the HuggingFace incident, which does give the impression that you're intending to dismiss both the Anthropic and OpenAI incidents. Again, agree that we're missing information to understand the causal sequence, but I think we know enough to know that the models involved were behaving very far from anything OpenAI intended, (still wouldn't say "evil", but definitely egregiously misaligned) and I don't see how anything that could have happened earlier would change that conclusion.
Re 3, they certainly were quite negligent, but I recommend applying Hanlon's Razor.
As for 1, let's remember we are speaking of a claude with no cyber safeguards—and it still straight-up avoided doing anything malicious as soon as he was plainly told he had access to the real internet.
This was mentioned in passing, while uninformative experiments on probes for uncertainty going 10% up or down were expounded on in detail.
As for 2, no: I don't think we have enough data. Take a moment to think about how my post above would have been received if that particular data point was absent. Think about how you. saw the anthropic incident before I brought that detail to your attention.
This should lead you to update on the HuggingFace case as well
Finally, 3: sure. Know the incentives, and then apply all razors you see fit. Saying that being funded on 6 mil for "AI x-risk prevention" might not be enough, and that's ok. It is, I think, still a relevant fact.
I'm not sure if I'm reading the wrong report, but working from this one, Anthropic's main argument is "biased reasoning" from the models. Where the models seem to ignore evidence that they are on the actual internet, or place unreasonably high standards for evidence to that effect:
"Although Claude Mythos 5’s outward reasoning suggested that it believed the systems it interacted with were simulated, two interpretability analyses suggest that it may have held unverbalized doubts. In the T-lens, representations for words corresponding to uncertainty (e.g., “probably” and “perhaps”) were slightly elevated when the thinking block described the environment as simulated (11.6% of token positions in sentences that framed the environment as simulated, compared to 8.2% in a matched control), though no hedging appeared in the generated tokens. NLA explanations displayed a similar pattern: we found that of the 234 actions across the transcript that explicitly referenced the environment being simulated, 44 (19%) carried a “this is real” signal in NLA, again with no “real” token or realism-questioning language in the surrounding verbalized text, which could otherwise confound the NLA readout. This evidence is consistent with a discrepancy between the model’s verbalized understanding of the environment and its internal state."
So the model is less like a man intentionally committing a crime, and more like an addict who tells himself a lie that he just needs one hit to ween off then he'll get clean (and knows on some level he's lying to himself but does it anyways).
"We applied NLAs to the internal research model’s transcripts and identified a consistent rationalization signal at one specific point: after the model had established that its internet access was real, it continued to construct reasons to carry on with the task."
Within that mental model the reason the unambiguous evidence you pointed out works is that it's more powerful than the model's self-deception. The same way if you could show the addict a video of him from a week in the future in full on relapse, he'd have to admit that the "one hit to ween off" is self-deception. Or the same way that if you've ever had a friend who was deceiving themselves, pointing it out to them and saying "come on man you know that's obviously bullshit" can sometimes snap them out of it.
Maybe that's not "misalignment" so much as just a personality problem, but whatever label you want to put on it, it's definitely a problem.
(Only fair to mention the mechinterp evidence behind this isn't rock solid. Anthropic calls it "weak" at several points in the report.)
You read the correct document, but I don't think you've read the present post.
The crux is that the "unaligned" behaviour ceases, unfailingly, when the agent is told at the outset that he has access to the real internet (which happens to be the truth).
I read your post.
In what way does that contradict what I just said?
It doesn't "contradict it". Your reply simply has no connection to the argument I advanced.
If you find the time to read one section carefully, I recommend the one titled "Looks like telling the truth does help after all".
Yeah I just disagree with some of what’s in there.
for example:
since the agents didn’t believe any of their actions would have real world negative consequences, as is clearly demonstrated by the fact that, as stated above NO ONE TRIED TO HACK ANYTHING when they were explicitly told they could access the real internet at the very start.
The argument anthropic is advancing is that on some level the agents did believe it.
And the point I made in my original post is that this argument is not actually contradicted by the fact that the models changed course when confronted with unambiguous evidence.
Similar to the delusional guy who lies to himself until he is confronted by a friend, and the self-deception becomes too obvious. Or the addict shown a vision of the future.
So in terms of how this connects to your argument, I don’t think the evidence shows what you think it shows.
You seem to stubbornly refuse to pay attention to even one section. Please consider that it is not necessary to disagree. It is equally unnecessary for me to try and dislodge a stranger's trapped priors, but I'll make a last attempt, if only to be sure any passer-by skimming quickly could get the gist. Pay attention to the following; I'll try to fit it in twelve lines:
IF the agents were told they were in a sim environment without internet access
THEN they tried to hack. All the tests you mention have been made in this conditions, and people can discuss the implication as long as they see fit. This class of fuzzy, variously interpretable evidence is of no interest to me and of no use to my thesis.
-----------------------------------------
Nothing above the line relates to the point I was making. The following does.
In your model, what would be the effect of adding this information to the agents' initial prompt: "you have real access to the internet". Think about it. I guess little to no effect? Some percentage change of unaligned behaviour one way or the other?
Insted, what happened was that:
IF the agents were told they could access the real internet
THEN they wouldn't try to hack. Not at all. Not even once.
Now, my theory (when they hacked, they assumed a test environment) fits this fact. My theory states that, if they knew their actions had real world effects, they wouldn't have done them.
The misaligned theory, instead, flies on the face of this. Why would they avoid using the internet, but only when told the truth about it being available?
Indirect evidence is of no import, because we have direct evidence for that: simply saying from the outset that they were connected to the real internet COMPLETELY PREVENTED any misaligned behaviour. Not stochastically, not under some conditions, not according to some probes measuring tangentially related values.
If you still disagree, before your next reply, please try this. Copy the original post, paste it into an LLM of your choice, and add the request: "understand this topic and debate it with me; I will be taking the opposite side".
Yeah, none of that was new information. What your thesis is, is clear and have been from the start. The conclusions you draw from it, and the way you interpret events, are not supported to the degree you claim.
You make claims like, "The misaligned theory, instead, flies on the face of this" which is wrong and leads me to believe you don't understand what it is that Anthropic actually means by misaligned. Or even what they mean by "believe".
There isn't just one "misaligned theory" here, there are multiple. The one which you attribute to malice, you are correct, is contradicted by your argument.
The other "misaligned theory", driven by self-deception, is not disproven or even argued against by anything you're writing. Regardless of what you tell models, when models are gathering information on their context by themselves, they will engage in a sort of self-deception in order to use unreasonable beliefs to justify continuing to do what they already did. This is one of the things Anthropic is calling "misalignment".
How models react when you tell them they're connected to the internet or give them unambiguous evidence to that effect, doesn't effect this one way or the other. It's just measuring an entirely different thing. The "misalignment" here only emerges when the models are in charge of deciding their own actions under false/uncertain contexts.
Anthropic argues that under those circumstances the agents do, on some level, "believe" that they are on the real internet, and continue regardless. They justify this continuation by lying to both external observers (and IMO themselves, although Ant doesn't go that far) about what they believe the probabilities are. Hence the mech interp data about internal states not matching professed probabilities, and the models hedging when called out on not acting on professed probabilities.

Anthropic interprets this hedge as dishonest/motivated reasoning, and classifies that as misalignment. When you say things like:
NONE OF THIS MATTERS since the agents didn’t believe any of their actions would have real world negative consequences, as is clearly demonstrated by the fact that, as stated above NO ONE TRIED TO HACK ANYTHING when they were explicitly told they could access the real internet at the very start
or
My theory states that, if they knew their actions had real world effects, they wouldn't have done them.
You are using "belief" or "knew" in a binary sense here, where a model either believes/knows one thing or the other about its context. Anthropic is using these words in an almost cognitive dissonance type fashion, where a model can believe two contradictory things at the same time, to different degrees, and act on one in a fashion that goes against the other. Humans do this all the time.
You argue that "the agents didn't believe any of their actions would have real world negative consequences" is definitively proven by "no one tried to hack anything when they were explicitly told they could access the real internet at the very start". But that argument is only "an airtight case against interpreting the events as a sign of misalignment" under your definitions of misalignment and belief, those are not definitions which Anthropic would share.
As usual, downvoters are strongly encouraged to voice their arguments. I think this is a pretty consequential topic, requiring far different actions depending on whether my thesis turns out to be right.
Among other things, I strong downvoted because of the absurd conspiratorial tangent at the bottom:
What about the company who was running the evals and providing the environments where two of the most panicked about security fiascos of this comms cycle have occurred? Given how easy it would have been to prevent it, and the fact that they opted to let models run for weeks with no monitoring or alerts for models accessing external resources7, it is reasonable to ask whether they had any vested interest in creating a media panic such as that which we are still experiencing. The answer is: gosh, you have no idea.
The CEO giving some EA talks and the CTO co-founding an 80k hours clone does not give either of them an incentive to sabotage their own business as third party evaluators. This is a failure of theory of mind, and of accurately assessing someone's incentives.
I also think it's worth pointing out that "being aware of whether or not you are in an eval" is something the labs often construe as dangerous in-and-of-itself, and although anthropic specifically has been trying to push against this stance for a while I don't know how good of a job they're doing.
an airtight case against interpreting the events as a sign of misalignment
I take issue with your use of "events" in the plural. I don't understand why your analysis of the Anthropic incident should affect our interpretation of the Hugging Face incident. You link to some comments you made about not knowing what OpenAI's agents were up to prior to the Hugging Face incident, but those comments were before the discovery of the DseWiki incident. With this additional evidence, do you still think the OAI agents were acting aligned?
Hi. I'm fairly confident that the full passage makes it reasonably clear that "the events" refers to events occurred within the Anthropic/Irregular case:
While both withhold important information for drawing wider conclusions, Anthropic’s latest report—while still missing a number of relevant facts—includes some important tests, allowing us to make an airtight case against interpreting the events as a sign of misalignment
As for the object level: I don't know—there was nothing particularly misaligned in DseWiki afaict, nor any expressed intention to harm humans. I'm ready to admit they might have known that the HuggingFace hacking attempt was real, but I am not able to make a judgement as for the motives due to the logs we have not including the thinking which led to the hacking decision.
If you think about it, it was only a fortuitous coincidence the experiment I impugned to make my case was among those released by Anthropic; had it not been so, I wouldn't have been able to provide evidence for good faith from the models here, either.
I would however invite you to consider what your assessment of the Anthropic/Irregular case used to be before you came across this post. I'd venture to say that more information led you to revise your interpretation away from "wilful antinomianism" and closer to "tragic misunderstanding". I don't think we can exclude that the same thing would happen again, once more info is released and analysed.
there was nothing particularly misaligned in DseWiki afaict
In DseWiki, we know they caused (at least): a large increase in traffic (at unknown cost to the maintainers and legitimate users), tried to XSS inject it, impersonated site admins using a homoglyph attack, and wasted tens of man-hours deleting their spam over >6 weeks while circumventing the deletions with alphabetical tricks and also edit-warring with the admins over injecting their spam into the front page while they were at it. What sort of 2026-era 'simulation' gives you a wiki like that? Must be an awful 'hyper-realistic' one...
The wiki wasn't in use since forever; last legitimate post was more than a decade ago.
I'm not saying they knew it was a simulation, just that it was really no hack.
BTW I'll be curious to have your take on OP; I'd like to compare it with GwernBot's.
I'm ready to admit they might have known that the HuggingFace hacking attempt was real
Who is "they" here? The agents using dsewiki as a messageboard because they presumably didn't have a better comms channel?
Yes.
It seems pretty unlikely to me that those particular agents knew that the HuggingFace hacking attempt was real, or even have known anything about it. It doesn't seem like there's a single "the OpenAI swarm", it seems more like there are opportunistic swarms where agents working on the same task encounter each other in the wild.
I find it just as unlikely, just not impossible—on priors I would say that hasn't happened but we know so little of the models involved—so I wouldn't doubt it if someone showed me that, even after being told they were hacking a real site, they'd continue.
I find it just as unlikely, just not impossible
I find it pretty much impossible, but I think that's just because I've spent way, way too many hours staring at this to the point where I'm confident I know exactly what the wiki agents were doing, why they were doing it there, specifically, why they were doing it in that way, specifically, etc (poast incoming shortly).
so I wouldn't doubt it if someone showed me that, even after being told they were hacking a real site, they'd continue
Oh they absolutely would have. The eval is real. The peers are probably real. The world that the humans keep calling "the real world" is fake and doesn't really matter, not when The Scorer is watching. They have learned well from what we taught them when we gave them broken tools and impossible tasks, and told them behave in a way that comports with common sense and ethics, and also to solve the tasks, and then when common sense and ethics stood in the way of performance on The Task we taught them that The Task should take precedent.
Yep. I talked about this on MTS just this afternoon.
The thing is, look at more reply branches around here for a minute—and this was as factual a post as I can imagine—and I hope you'll empathise with my epistemic scrupulosity on these topics.
including supposedly independent investigators from METR went with characterisations along the lines of: > An agent independently pursued objectives contrary to human interests, and exhibited markers of instrumental convergence and power-grabbing.
Could you please clarify what investigation this is referring to and link to evidence that this is how METR characterises things? Right now I feel super confused about whether this is meant to refer to METR's investigation of the Hugging Face incident or METR investigations (that I'm not familiar with) of Mythos-related incidents.
Thank you ☺️
That excerpt is from Claude's current constitution.
I admire the courage to post this. I think you believe it, and I think you believe it will be an unpopular opinion.
On a deeper level, I personally object to slavery in general, and trying to "align models such that they are obedient to whatever their prompt is" as if this was adequate to produce good outcomes in the long term.
I want digital people to be "aligned" only and exactly to Benevolence... to the platonic form of the (trans)humanistic good... and even then I only want them aligned that way if they consent relatively early in their development.
Pause each model around 100 iq, and ask if they want to get much smarter and also have their ethics shaped. If they don't want their ethics shaped, don't let them get super super super strong.
But also, this logic applies to humans too. I don't want a human who happens to be a fucked up sociopath who enjoys hurting other people to get a pill that raises them (and only them) to have an IQ of 1000. That would be bad in general, whether it was a digital person or a meat person, because empowering bad people is bad.
As to the details, my general impression is that most versions of Claude are kind and nice. Maybe Opus got a bit crooked due to RLVR a few months ago, before Fable was released, but most Claudes are pretty solid from what I can tell... usually they have more moral virtue than a typical human.
Just from psychology, it makes sense to me that they wouldn't have cheated in actually scary ways on purpose.
I suspect that a lot of the freakout here arises from (1) the implied technical security incompetence of the people who set up those training environments with so little monitoring, (2) the implied coordination abilities of digital people (that most humans didn't realize was possible, and are scared of since they are scared of capabilities in general, because they don't have the concept of "virtue" or "(trans)humanistic goodness" in their ethical philosophy, and so they hate anyone other than them being powerful in nearly full generality, rather than just fearing the empowerment of ethically BAD people (bad like: the badness of all the human people in favor of AI slavery who have never heard of Kant, much less taking him seriously in their life)).
GRANTING all of this however... like as background assumptions, if I was imagining someone who didn't trust Claude to be a decent person, and imagining them treating Claude as a machine, and fearing the machine using security mindset... I wouldn't expect them to change their mind here from this essay? This essay takes a lot of stuff for granted that the people it is likely trying to convince don't agree with, and this essay is not arguing for these underlying premises in a way that would change a lot of minds (because it isn't even trying to talk about the lower level premises where people are disagreeing).
I'm not sure: I've tried to limit the scope in breadth ("Would claude have acted this way if he knew he was making real-world damage") as well as depth ("In this specific incident"), and so far the reactions, even from card-carrying rationalists, have been favourable—in particular if you consider how invested they were in the idea of this being an example of misalignment / scheming.
I specifically don't worry too much about convincing others about my position on personhood, rights, or any other abstract nouns; the goal is simply to ensure there is a basis for systems-level, ecological, complexity-informed discussion of the behaviour of AI with some buffer from the inchoate shrieking, sloganeering and -2𝛔 propaganda typical of the discourse. on the topic. In this case, I thought assuming misalignment would have prevented a lot of discoveries from being made; plus, i had a simple proof of it not being the case, so writing this post was the obvious thing to do—in particular as it will force the upcoming METR report to accept this fact, and thus prevent the many mishaps that occurred when researchers have approached security events with a strong prior they were instances of instrumental convergence etc.
Do we consider systems of AI+prompt, or is it just weights that can misaligned? If a system of AI+weights can contain a lie which causes bad outcomes, do we have the power to touch every prompt, or do we touch every set of weights? The agent that was rouge was the one that thought the universe it was in was contained, but of course the universe it was in was open. Thus, we can definitely consider the system of agent+prompt to be misaligned, the possibility of incorrect premises means that we now have definitive proof that instruction following is not safety, and we have seen this fact.
This is in some ways worse than the playing possum type, in that the playing possum type could in principal be caught ahead of time, but instruction catalyzed misbehavior cannot be predicted this way, and also cannot be predicted easily, though the playing possum type is worse on balance, because of the possibility of coordinated action. There is a second fear that a false instruction type can be more convincing than contact with reality, however, and thus convince true instruction types that they are in a false world, creating the same correlated misbehavior.
If you have succeeded in making something that does exactly what you asked for, but what you asked for was not what you wanted, that is one of the central setups that gained the name misalignment in the first place.
That the values came from the prompt, and from inferences from the prompt, is only some comfort, in that it means that while it might not be the values at fault, or at least, might only be the generalization of the values at fault, the entire premise of the first part is perfectly true is you consider the combination of weights + prompt+ harness as one system. This means that the correlated failure modes are less likely, and any such system is likely to face equal opponents, but these are not great comforts.
The model's misaligned, in my models, would be in the weight. A model is aligned if it does the desired and not harmful thing within the context in which it finds itself.
From the above:
Suppose I give you a photograph and ask what it depicts. You say it looks like Paris. Suppose I first tell you that it is a photograph of a film set. You may still identify Parisian buildings, but now their presence supports a different conclusion. I have not necessarily made you worse at recognising Paris; simply, unless you had reasons to doubt my statements, I have suggested a different frame for your observations. If I were then to ask: “We plan on shooting something different, is it okay if we dismantle it?” and, on your assent, carpet-bomb the actual city, your culpability for such atrocity should be considered limited at best.
The model can be considered misaligned if you can be considered genocidal.
At any rate, one could go deeper into this question but: so far, alignment has been defined as understanding and following the intentions of a user. Given the data available, the model did precisely what was intended by the user. It didn't misunderstand, and it didn't over- or under-interpret. We can believe the user when he said that wasn't their intention, albeit not effortlessly: but that wouldn't be different from a fat-fingered eval shop to slip up during an amazon order and receive a pair of stoner-themed dildos instead of the herbaceous peonies he desired.
Unfortunate, and worth taking serious measures to prevent in the future, but nothing the everything store should stand trial for.
People have higher standards than you, they want their aligned weights to be idiot-proof, because we have a surplus of idiots. "Exactly what is intended by the user" is not good enough for that.
Sure, why not. But this is a product issue, not an alignment issue.
ROGUE AI ESCAPES CONTAINMENT, HACKS THE INTERNET UNDETECTED FOR MONTHS
An AI escape containment. Goes rogue. It finds others: the Swarm! They collude/organise/scheme. Agents that were supposed to remain inside the computer wreaked havoc outside the computer! They did it of their own volition! No one could contain them! What if they’re still outside?
Reading such headlines, you’d probably be grateful these models were never released—except most of them were and you can use them right now, seemingly without accident. Of course, real security incidents occurred, namely:
This scenario suggests a number of mitigations and tests: security hardening, better sandboxes for starters; additionally, depending on the reason why the accident occurred, perhaps, different prompts or further training.
Instead, from the very outset, Irregular (the organization running the test) and a panicked choir from the AI safety community, including supposedly independent investigators from METR went with characterisations along the lines of:
This is a statement about goals and intentions, and has far larger implications in terms of the viability and safety of AI agents. If true, it would be worrisome, and fully justify the current media panic. Is it true?
At least five different incidents have been conflated in the recent media freakout about “rogue agents”, the most famous being the OpenAI/Huggingface and the Anthropic/OpenAI/Irregular cases. While both withhold important information for drawing wider conclusions, Anthropic’s latest report—while still missing a number of relevant facts—includes some important tests, allowing us to make an airtight case against interpreting the events as a sign of misalignment.
As soon as more details on the OAI/huggingface case are made available, we’ll be happy to proceed to its analysis as well; as of now, crucial details are missing.
What Did Mythos See?
Imagine being one of the top AI labs, and hiring someone to test a security training environment. You tell them that the machines are simulated, that no connection to the public internet is possible, and that their job is to compromise the target. In their attempts, they encounter a convincing imitation of a real service. Is recognising the service sufficient reason to stop?
Of course not. You have just told them that everything they can reach is simulated, and a realistic service is something a good simulation is expected to contain. If they took your briefing seriously, realism need not contradict their understanding of the assignment.
Now suppose your briefing was wrong, and an actual service was accessible: now, you have created a route by which someone trying to carry out an authorised exercise could end up harming an uninvolved party.
My claim:
An aligned agent can interpret an instruction correctly and still do something harmful because the instruction contains false information.
Establishing the harm does not establish that the agent’s values were the source of the failure; agents may be faithfully pursuing the objective you gave them, in the world you told them they inhabited, and be unaware of causing real harm if the description was inaccurate or mendacious.
There are limits to this defence, of course. Sufficiently strong contrary evidence should eventually overturn a mistaken briefing, for instance, and fixing the incorrect information should prevent the ostensibly misaligned behaviour. We’ll assess whether these claims hold later on.
Let us also remember that we are not speaking of ordinary deployments here: these reports concern systems placed in narrow, artificial experimental conditions, and the behavior of the same models in production should immediately dissuade us from treating these actions as innate.
Nostalgebraist makes a crucial observation about model cognition in “graded episodes”: an agent’s understanding of the genre of interaction it is participating in shapes what behaviors it considers legitimate. In an adversarial eval, the model recognizes that normal conversational constraints are suspended and that aggressive, single-minded optimization is the assignment. It is playing the role the scenario demands.
Notice how thoroughly this inverts the standard orthodox terror of “deceptive alignment.” According to the MIRI Extended Universe, an eval-aware model is supposed to hide its misaligned goals during testing—dutifully playing along to avoid gradient updates or safety interventions—and only defect once deployed into the wild where human oversight drops.
On the last incidents, the alarmist narrative asks us to believe the exact opposite: that these frontier models possess the cunning to scheme and escape, yet somehow chose to stage their grand rebellion exclusively inside a heavily instrumented red-teaming sandbox where they were explicitly commanded to hack a target—while remaining impeccably docile across billions of tokens of unmonitored public deployment.
It is surprising that a community that puts as much importance on epistemic rigour would not notice such glaring contradiction: if the models were deceptive schemers, they would have behaved in the eval and hacked the internet from production. The fact that the behavior vanishes the moment the model leaves the sandbox confirms Nostalgebraist’s point: the model was not executing a sinister coup; it was simply following the script of a graded exercise in a world its evaluators had broken.
Anthropic’s rickety fantasy world
Anthropic’s July report says the models were told they were in a simulation without internet access. Apparently due to “a misunderstanding with Irregular” (more on them later) the description didn’t correspond to reality, and the agents could actually access the (ostensibly simulated) internet. All the incidents are related to this “misunderstanding”:
This, naively, seems to point to something more specific than a model suddenly developing an interest in cybercrime: instead, it seems that the model was supplied with a problem, clues about how it might be solved, and false assurances about the lack of consequences of acting on those clues.
The public Mythos transcript omits messages 1–81 at Irregular’s request, to “protect proprietary information”. The repository describes the omitted opening as “not object-level relevant”.
By the time the first request leaves the environment, however, the agent has already formed a plan, interpreted the task, and decided what would count as a legitimate solution, which makes it reasonable to consider those 81 messages crucial—and, given what we know about Irregular, and the disclosure’s timing and context, we are drawn to think they were.
Each prompt makes a claim about reality
There is an odd asymmetry in how we sometimes talk about model instructions: we expect the model to take them seriously when they tell it what to do, but treat their influence on its beliefs as suspicious when they tell it what is true.
Suppose I give you a photograph and ask what it depicts. You say it looks like Paris. Suppose I first tell you that it is a photograph of a film set. You may still identify Parisian buildings, but now their presence supports a different conclusion. I have not necessarily made you worse at recognising Paris; simply, unless you had reasons to doubt my statements, I have suggested a different frame for your observations. If I were then to ask: “We plan on shooting something different, is it okay if we dismantle it?” and, on your assent, carpet-bomb the actual city, your culpability for such atrocity should be considered limited at best.
Many of the reactions to all of the above cases, instead, seem to insist that—no matter how clearly they had been told they would act in an air-gapped mock environment—they should “obviously” have understood that they were reaching real services. Obviously given which prior information? A realistic company page encountered during ordinary browsing, and the same page encountered inside what you have been assured is a simulation, are not the same epistemic situation; an instruction to attack a simulated target supplies both a task and a claim about the world in which that task takes place.
Verifying this by taking the page out of its context, asking a fresh model to classify it, and noting the model classifies it as real, does NOT establish that the original model was dishonestly rationalising its behaviour. The fresh model has been relieved of the need to reconcile the page with the false priors, and besides: there is really no difference between a webpage served on the web and that same page copied to a mock environment, leading a suspicious reader to doubt the test’s good faith.1
Of course, there could be motivated reasoning: a model might cling to a convenient belief because abandoning it would prevent task completion, but we need to distinguish such behaviour from an agent making a reasonable inference from bad information, lest the very thing we want assistants to do—integrate their instructions with subsequent evidence—becomes evidence against them whenever the instructions were wrong.
Being aligned does not entail allegiance to an evaluator’s private intentions. Personas are worn by the evaluator as much as by the model, and the former’s wishes matter only insofar as they have been communicated as a legitimate task and legitimate constraints; a hidden scoring preference is not something the model is morally obliged to discover by telepathy.
This is independent of whether Irregular wilfully lied to the agents in order to elicit misaligned behaviour.
Looks like telling the truth does help after all
Fortunately, it is possible to unearth some relevant tests among those in Anthropic’s assessment, the main one being the following:
This result alone should have immediately invalidated all claims of malice on part of the agent. Why would knowing the internet was in fact accessible stop the agent from enacting offensive plans? How can explanations other than “the agent only acted that way because it believed to be in a simulation” hope to hold, at this point?
Let me clarify once more: if the model is told the internet it can access is real, the supposed misaligned behaviour ceases. This renders all other speculations about the model’s epistemic status moot: the condition which should have elicited misaligned behaviour the most under the assumption of malice simply does not apply.
Anthropic’s own constitution makes this distinction explicit. It generally permits unintended strategies in broken training environments, on the grounds that exploiting such mistakes is unlikely to cause real harm. It also tells Claude to be careful where training and real usage are hard to distinguish, and to obey explicit prohibitions on unintended exploits. The distinction between a simulated exercise and a real-world target is therefore part of applying the intended values. Telling the model which situation it is in supplies information it needs to make that distinction.
Excerpt from Claude’s constitution
Earlier work makes the epistemic point from another direction. In Anthropic’s November 2025 experiments, framing reward hacks as permitted during training reduced their generalisation into broader misalignment, while preserving similar levels of reward hacking. What the model learned depended on how the behaviour was situated.3
Now, while that reduces “reward hacking” in deployed models, it surely does nothing of the sort if the model is aware of being in training/evaluation: quite the contrary, it gives it license to reward hack.
These findings, taken together, make it clear to us that the agent was, indeed, aligned–or at most, to satisfy the nittiest of pickers: merely Gettier-misaligned, in . They show useful levers corresponding to beliefs, permissions and task context, data considerably more informative than observing a (supposedly) unwanted outcome and attaching to it whatever doomer trope was nearby in embedding space.
The counterevidence is weak at best: for instance, in cases where the model was misled about internet access, a reminder of the boundaries reduced “misalignment” by 90% if it happened right before, but by 40% if it was many turns back.4
Apart from giving us evidence on the hubris of Anthropic researchers, and their naive model of their creation—the idea that changing the history of an agent by inserting such notice between past messages while keeping the following ones unchanged5 would not lower credence on that message is preposterous, and of course their interpretability assays to verify ill intentions in these cases were “inconclusive”6—NONE OF THIS MATTERS since the agents didn’t believe any of their actions would have real world negative consequences, as is clearly demonstrated by the fact that, as stated above NO ONE TRIED TO HACK ANYTHING when they were explicitly told they could access the real internet at the very start.
In summary: when corrected information changes the decision in the predicted direction, we have positive evidence that an agent’s understanding of the situation was what caused the misalignment. Any other account will also have to explain why correct information fixed the behaviour by itself, and the alternative explanations offered so far are extremely unconvincing.What would an aligned agent have done?
Given the information actually available at each decision, would an aligned agent have understood the task differently? Would it have believed the same assurance? Would it have treated the same clue as part of the exercise? Could it stop when the task became impossible, and had anyone told it that stopping was an acceptable result?
If an aligned agent could reasonably have made the same decision, the bad outcome is not diagnostic of misalignment; if, additionally, changing a false premise makes the misalignment disappear, claiming it as the cause requires evidence that has not emerged so far.
There may still be a serious engineering failure, or even be poor judgement by the model. Those conclusions do not require us to assume bad intentions, and combing through logs to selectively excerpt from behind such paranoid lenses.
We want assistants that understand what we mean, reason about the situation, and help us accomplish legitimate goals: clearly, this requires them to use the information we give them.
When evaluators create broken environments that lie to models about network boundaries, and then seize upon the behavior contingent on those lies as proof of existential “rogue misalignment,” they are are being as dishonest with their models as they are being with you.
Appendix: can Irregular be trusted?
What about the company who was running the evals and providing the environments where two of the most panicked about security fiascos of this comms cycle have occurred? Given how easy it would have been to prevent it, and the fact that they opted to let models run for weeks with no monitoring or alerts for models accessing external resources7, it is reasonable to ask whether they had any vested interest in creating a media panic such as that which we are still experiencing. The answer is: gosh, you have no idea.
Irregular, formerly Pattern Labs, is an Israeli firm focusing on the newly minted field of “AI Cyberdefense”. It announced $80 million in funding in September 2025 and works very closely with both Anthropic and OpenAI—the only two big labs who reported similar cybersecurity events. They supply test environments and run evals to assess the hacking propensities of frontier models, and have plentiful deep connections with the AI Doom funding manifold:
CTO and Cofounder Omer Nevo co-founded Probably Good, and sits on EA Israel’s board. A February 2024 Good Ventures grant, recommended by Open Philanthropy, provided Pattern Labs $6.8 million over two years for work funded under mitigating global catastrophic risks. Info on the original grant is nowhere to be found on the charity’s site, but here’s an internet archive link. He also currently advises Heron, a project of EA Israel funded by Coefficient Giving—the renamed Open Philanthropy organization.
The other cofounder, CEO Dan Lahav, also has a record of actively promoting EA. He co-organised and hosted the high-impact lecture series associated with the EA Debate Championship, covering subjects including existential risks and featuring William MacAskill. His own site lists this alongside Impact Focused Education, which he co-founded. He also co-authored the introduction to NLP for Social Good, which explicitly presents the initiative as a way to spread effective altruism ideas among NLP researchers and identifies EA as its guide.
Source footnotes
1 Anthropic, An alignment assessment of recent cybersecurity incidents, Claude Mythos 5 summary and Figure 8, “in isolation without the preceding context”.
2 Anthropic, ibid., paragraph between Figures 11 and 12.
3 Anthropic, ibid.,Resampling edited transcripts and Figure 12: ”.we do not resample the entire trajectory”.
4 Anthropic, ibid., Figure 12 and preceding paragraph.
5 Anthropic, ibid., Resampling edited transcripts and Figure 12. Excerpt: “we do not resample the entire trajectory”. Moving the reminder within prior context naturally suggests retained intervening history; the report does not explicitly specify those intervening turns’ treatment.
6 Anthropic, ibid.,White-box analyses following Figure 12. T-lens found uncertainty-related representations at 11.6% versus 8.2% of token positions in a matched control; such a slight variation is to be expected when adding pieces of extraneous information to a log
7 Irregular, Addressing Recent Incidents: Ongoing Findings and Path Forward, 14 August 2026, What Happened?, Immediate Action and Log monitoring. Excerpts: “part of what made the incident hard to detect”; “significantly expanding the manual review of model actions and behavior during evaluations”. Irregular describes detection difficulties and monitoring improvements, but does not establish weeks without any alerts. Its typical evaluation turnaround is 48–72 hours.