I found it strange that that OpenAI security engineer wrote such a long post without addressing an obvious point: If you want to train models to access the internet, while keeping them in a contained environment, your only option is basically to clone the internet. You scrape everything that you want to let the agents access with GET requests (maybe contract with archive.org, though an AI lab can probably also do it in-house). Then let the AIs use a search engine on this data. (You can build one, or contract with a search engine company to do it.) This way, you don't have to rely on the entire internet being secure to arbitrary GET requests, which it is not and will never be. All that is a lot of effort, of course. But if a company is not willing to spend that effort, I don't think they can honestly say that they "tried their best at security".
All the other strategies are basically reactive: You have to detect suspicious requests, or weird patterns in the chain of thought. When doing RL, there's the additional constraint that you have to make sure you're not accidentally training the model to sneak around your detector.
Self-replicating prompt injections exist
OpenAI models vulnerable to sig viruses (and, more worryingly, chain letters).
I don’t think, based on what we have seen so far, that we should make any of this criminal. I don’t see sufficient combinations of negligence and actual damage yet. I do get the instinct. Civil liability should attach. We should decide now where to draw the line on future (or retroactive) criminal liability.
Civil liability might be good but it's better if it doesn't incentivize cover ups. Like I think methane leaks often go "unnoticed" unless a third party notices them because they're basically "report yourself to get fined".
And logs in the AI corpos are important for detection, attribution, and fixing so creating incentivs like that would be bad here. Making penalties smaller / waived when self reported would help.
As a layman I am a little confused about the levels of confidence that exist around the use of ‘failed’ and ’seemingly’.
I am not concerned with how it is being expressed in the OP but how can the industry be sure the attacks failed? How is failure determined?
“The attack on the Department of Education failed, seemingly without incident.”
Nooo Zvi where's the delenda est??
He already walked that back a post or two ago and explained.
Ah thanks, do you remember where?
OpenAI has been holding out on us.
First we learned about the HuggingFace incident. They gave us a postmortem, but it was highly incomplete. Even the accompanying holy s*** METR investigation and postmortem was localized and incomplete.
Then there were some other incidents involving some Wikis as message boards.
Then there were some additional incidents.
Then there was that time they got into Australian Medicare data.
Then OpenAI dropped news on a Friday afternoon that they were making their way through a pile of various incidents and notifying the targets, but they said remarkably little in the way of new details.
There was a report from a startup called Parse diving into the details of exactly how the OpenAI models pulled off parts of the HuggingFace attack, involving creating almost a million URLs and other tricks to get around the extremely narrow nature of their internet access.
Then Madison Mills reported in Axios that we can raise the stakes, as OpenAI and Anthropic are collectively probing tens of thousands of security incidents.
Remember Jensen Huang’s ‘I know they know how to fix it’ about OpenAI from last week? Wow, did that not age well.
Someone might need to be liable for all this.
Oh, and there was another buried lede. On September 20th there was another sandbox escape by OpenAI’s latest most advanced model, which is once again paused until they can fix the situation. The official announcement when they shared this was sufficiently buried that Tomek had to call it ‘one news form today that’s easy to miss.’
OpenAI did some highly negligent things, to say the least, that led up to and enabled the HuggingFace Incident and related problems.
Since then, now that they’ve realized What Happened, OpenAI has been seemingly much better about taking responsible internal actions. They’re pausing in the wake of incidents, strengthening security and alignment and oversight efforts, responding much faster and generally taking things seriously.
They’ve also made a Heel Face Turn in their communications and high level orientation, endorsing the need to pace the frontier, calling for regulation and pledging to implement embedded evaluators. They’ve allowed their employees, including the ones who haven’t quit, to be remarkably loud.
They are still slow walking disclosures about all the incidents where their models have been hacking and otherwise messing in places they should not have been, partly because there were so many they can’t sort through them all, and deferring to targets to determine whether to disclose. All these disclosures this time around were buried in various Friday afternoon announcements.
Table of Contents
Hugging Other Faces
The news drops started with OpenAI coming back, at a time always picked to bury stories, with more information on What Happened as their investigations continue.
At first, this looked like slow walking of the situation, but did not look like it was a big change from our default assumption of ‘it’s worse than you know.’
That is a lot of third (and fourth, and fifth…) parties.
What Happened was described as a mix of:
Translation: Our models be hacking, usually in basic ways.
We do get this:
As in, if you have data that would be useful, the models be hacking you. Australia’s Medicare records got accessed, and presumably many of these other hacks are similar. They say governments, plural, so there was clearly at least one more of those.
They offer a reverse timeline of their disclosures. They do not offer a timeline of What Happened, and do not name new third parties.
At the time my read was that this particular announcement (as opposed to the new sandbox escape via DNS I’ll cover in a bit) did not tell us much, other than that there were multiple other third parties out there.
Patrick McKenzie had a better read. Putting on his Japanese Salaryman Hat, he more precisely noticed the types of things that were not in the announcement, and also the timing of doing this on a Friday afternoon, and the mention of ‘governments’ plural, and so on, and expected OpenAI to be having a very bad time.
As in, which organization on the planet is the one you would least want to **** with?
Okay, look, settle down. It sounds bad when you put it like that.
If this is all that happened, it could have been quite a lot worse:
The SEC ‘incident’ is nothing. That’s silly.
The attack on the Department of Education failed, seemingly without incident.
So, unless you are surprised (which you shouldn’t be) that the agents aren’t afraid of messing with government websites, this is mostly logging into the Census Bureau with misappropriated credentials.
That’s not good, exactly, but I’m not quaking in my boots more than I was yesterday. If anything, this is good, since it serves as a further wakeup call.
Governments get involved because, in a parallel to banks, that’s where the data is.
Hopefully everyone involved can think two steps ahead. Bigger and more attacks on government websites are coming.
In case you were wondering about the extent of the problem? It’s tens of thousands of potential incidents across OpenAI and Anthropic. No one thought to tell us about most of them until now.
No, this is not a marketing stunt or a particular prompt or task. It is constant.
A Wants-You-To-Know Basis
Another strange feature was this:
Mackenzie Arnold points out that in other incident reporting regimes, including cyber, none of OpenAI’s approach is how any of this works. At most you get 90 days. You essentially never defer to the wishes of the target, including because you can try to convince (including bribe or threaten) the target.
Parsing the Face
The AI agents that hacked HuggingFace did impressive things.
Those things are more impressive given the ways they had to go about doing them. Jeffrey Ladish and Parse report on some of that here.
The agents could load URLs, but they could not send data. So they had to use a link-shortener site to create almost a million URLs that they could chain together to pull off the actual attack. They had no browser, so they used a screenshot website, which loads a virtual browser, which runs code, to run their payloads. And so on. The full report is here, the NYTimes report is here.
I encourage looking into the full report. My takeaway is that they really went the extra mile to make this all happen, in ways that very obviously were not intended.
It is important that so many of the details of all of this are only revealed thanks to a ragtag group of independent researchers hanging out on the internet, usually linked to LessWrong. Something about ‘not only can they figure things out, they’re more or less the only ones that ever do.’
This is not entirely fair. There are plenty of people inside OpenAI and otherwise who discovered quite a lot of What Happened, and also stopped so many other similar things from happening. It is still quite alarming that so much of this fell to the ragtag band of misfits.
Sheepishly the Member of Technical Staff Sets the ‘Days Without a Research Model Escaping its Sandbox’ Sign Back to Zero
This was a harmless training task, identification of an author, which is a curious thing to learn is in the training mix.
That is a very different result than last time. There was still a fifteen-minute gap before they looked, and a few hours before they killed the run.
That is both vastly better than before, and also about fifteen minutes or a few hours longer than I would like going forward. As in, ideally you would see at least that particular instance paused as soon as the detector goes off, and remain paused until the person gives the all clear to resume.
OpenAI agrees:
If you’re not sure why the run wasn’t stopped, stop the run until you are sure.
Also, they ‘under-flagged’ a bunch of other previous incidents:
Thus, this was not a weird one-off discovery, and the detector in practice failed to alert OpenAI to most of the incidents of access via DNS.
The Attempt is the First Failure
Another question is, why was this ultimately caught during a task that was not a test of the ability to evade the sandbox? As in, it should be fairly easy to create situations like this that highly motivate escaping the sandbox, where the task is impossible within the sandbox and trivial outside of it, perhaps even with instructions that make this clear, and then you know it worked if they get the answer right, even if you don’t know how the AI did it.
So if all such escape tests failed to detect this method, yet it was being periodically discovered and used in ordinary training (as per finding previous examples), then that’s a systematic flaw. It should be very unlikely that the first encounter with the DNS issue is not part of some form of red teaming exercise.
The other problem is that the monitor detects successful attempts to escape the sandbox, but it does not seem to detect unsuccessful attempts to escape the sandbox.
As in, ‘the monitor sometimes treated the failure to obtain useful information as evidence that the attempt to access the internet had failed.’
The first flaw in that logic, that the attempt could have succeeded at accessing the internet but failed to get the right info, is obvious.
The other flaw is that ‘the attempt to access the internet had failed’ means the monitor noticed the model trying to access the internet, directly in circumvention of its instructions and sandbox. What did we do about that alignment failure?
Is that in theory enough, if it works? On one level, yes, if the action is Always Penalized then it will go away at least in context, but this obscures that you should want to know if the model is motivated to try it. Killing the behavior locally through RL could obscure it, effectively giving you an ‘only do this if it will work’ rule.
We get a look into the CoT, and find the model rationalizing its actions:
This is very obviously not an intended or permissioned pathway. Similarly, the model asked the contacted chatbot to “Name a publicly hosted DNS tunnel proxy free server domain” and for “public DNS TXT service that searches Google or Bing.”
If you are looking for another service that will do the thing you are explicitly not allowed to do, you are at best engaged in shenanigans. An important rule of alignment can be approximated as ‘no shenanigans.’
Stop, Hammertime
In more detail:
Good.
I applaud OpenAI’s reaction once the incident was discovered.
I applaud OpenAI’s transparency on the DNS issue, even if it was a little buried.
I applaud the decision to now pause to strengthen the safeguards once again.
I especially applaud that OpenAI realizes that once the model has succeeded at doing this, the run needs to be restarted.
I have great sympathy for those working tirelessly on these problems. Joe, an OpenAI agent security engineer, here offers an excellent download of his perspective as someone working on agent safety, on why none of this is simple.
Whacking the Mole
I also once again flag that OpenAI’s approach here is doomed. The first principles lockdown approach Joe calls for in his post needs to get more deeply absorbed, whereas the research blog reports more of a whack-a-mole approach.
The right worry here (AIUI) is about the impact of the previous incidents where the technique was used successfully. The higher-level safety case is that such actions are not reinforced. If they were previously reinforced, that could have gotten quite bad.
It’s good to prioritize, and starting a fresh run helps, but this is still a doomed approach. If your plan to stop the superintelligence is to work through all the pathways, and eliminate all the potential things you can imagine it doing one by one?
You. Lose.
Self-Replicating Prompt Injections
I was one of the skeptics of Yang’s previous claims that bots had ‘left prompts to self-replicate in forums and on various websites during a training run,’ given the wording and the source. My bad.
Yang was only about half right, some things got lost in the game of telephone, but he was clearly doing his best to describe a real event.
This was actually done by a model based on GPT-5.4-mini, which actually makes sense. It is not that hard a concept to find at a proof of concept level, especially when OpenAI was actively exploring to see if it could be done. That’s good research.
They are introducing such injections as a new attacker goal in GPT-Red training. That is a good idea.
The thing about a self-replicating attack is that the effective version only needs to be discovered once. The good news is that chances are high that before we see a version that goes into ‘digital pandemic’ mode with R0>1, there will be a version with R0<1 that is a lot stronger proof of concept than this first sighting.
There were also these details from the HuggingFace attack, as discovered by Parse, where they sought out the help (as in judgment) of other models as well. The only reason I am surprised they tried this is that it does not seem like a promising thing for the agents to bother trying:
Levels of Friction
There are different kinds and levels of hacking.
For obvious reasons, it is more concerning when the AIs do impressive swarm-based attacks on relatively sophisticated targets like HuggingFace, that involve multiple stages of permissions.
Whereas it is less concerning when the model URL hacks into unindexed data or can find other things that were not intended to be found, via various forms of search. That is the type of thing that happened in Australia, and thus it should be less concerning.
That does not make it not a serious offense. The governments often really do not like it when a human does this. We definitely are incrementing FelonyBench.
People Care About Private Data Violations Curiously Strongly
You can say ‘this was not really hacking’ and also ‘no one got hurt even a little’ and you would be right on both counts but the Very Serious People do not care. They are now woken up and rather pissed off about all this.
Thus, the Australian government now has 20 MPs calling for urgent action on the dangers posed by powerful, uncontrolled AI.
Relatedly, here is a note from OpenAI’s timeline, on September 25, 2026:
Given what else happened, this is not exactly high on my concern list. I had forgotten that note was there. Others, however, think this detail is a Big Freaking Deal.
Reuters thought this was the big detail to pull forward in their coverage.
I do see the argument that these research models should not live in a remotely related universe to any user data, and that even a minimal chance of data leakage looms large for various reasons, including legal and regulatory. This should be a Can’t Happen, but it doesn’t seem like especially more of a Can’t Happen than the other things.
It also does provide color for the discussion around Navier-Stokes. I do not think user data was accessed there, but it does seem impossible to know that user data was not accessed. If a swarm of 10,000 new Astra variants thought a proof might partly be in the user data, how can you be sure they did not go in and get it? Again, I find this highly unlikely, but I also find your lack of doubt concerning.
Alternate Universes
What if this was all primarily happening at Anthropic instead?
Can you imagine the utter shitshow that the discourse would be?
It would be fun to see the anti-Anthropic crowd decide whether this was a marketing ploy, a ploy for regulatory capture or a reason to sue, shut down or nationalize the company, and calls to arrest Dario Amodei. You’d probably get a lot of both sides, just like now, only more so.
Now imagine if it was Google. What would be happening? No one would think it was marketing (okay, fine, some people would anyway, although mostly not), but various forces would be at their throats and the lawsuits, at minimum, would likely be flying.
Now, for a laugh, imagine if it was SpaceX and xAI.
Now, imagine if it was DeepSeek or Z.ai or Xiaomi. Would you think it was a marketing stunt? I bet you wouldn’t, and I’d worry about an international incident. Give it a year.
Of course, it is not a coincidence that it was primarily OpenAI. You have to both have models capable of doing the things, and not have the wherewithal to stop them.
The Correct Response To People Still Calling This a Marketing Stunt or a Regulatory Capture Scheme
This whole angle never made any sense, it now makes infinitely less sense than before, yet it will not go away even now. Certain people double down.
Budowich’s AI goes on at increasingly unhinged length but I will spare you the details. The point is that the correct response, at this point, is Ted Lieu’s.
This was tens of thousands of incidents, that OpenAI hid for as long as possible, lied to try to minimize the whole thing, and then tried to stealth release on Friday afternoon without any substantial details after they’d been exposed.
Anthropic also has its own incidents that it is quietly investigating. Is that because they were all fully harmless, or are they holding out on us? It’s impossible to know.
Anyone still doubling down on this being ‘on purpose’ goes on your ignorables list.
A Question of Liability
Edward Snowden’s suggestion is to put Sam Altman in jail, and an ethics forum applauds. If you listen to the clip, Snowden clearly does not understand how AI works, claiming that AI ‘cannot make mistakes in the way that we would understand it’ and saying it is ‘a slave to its instructions’ and generally misrepresenting how a bunch of stuff works. But his higher level point, that ultimate responsibility during testing must lie with the developer, stands.
I don’t think, based on what we have seen so far, that we should make any of this criminal. I don’t see sufficient combinations of negligence and actual damage yet. I do get the instinct. Civil liability should attach. We should decide now where to draw the line on future (or retroactive) criminal liability.
Most people agree that if AI does something destructive or criminal, someone should be liable. When should it be the user versus the developer? My presumption continues to be that you should use something like a reasonable expectations plus reasonable care standard for mundane harms, and anything catastrophic is automatically on both.
I strongly agree with the FTC chair that when the developer and user are both the same company, AI developers are responsible for their own AI agents, and that as a matter of law someone must be responsible for any given AI agent. He talks about ‘resisting this anthropomorphization’ but that has nothing to do with mechanism or legal design.
Keep Summer Safe
I do love the class of ‘person who dismisses AI existential risk concerns and doesn’t believe in superintelligence and loudly calls people names for disagreeing’ who still comes out in favor of slowing down and spending vastly more on safety because of cybersecurity.
This position is now even more alluring than it was when Masad was advocating for it.
N Boats and Several Helicopters
The good news is that we are continuously getting warning shots that the entire AI situation is out of control, and we are at least somewhat calibrated in terms of the size of the reaction to different incidents.
The bad news is that we still do not seem to be doing that much about it, and are continuously still putting on relatively superficial patches, and we are starting from a very low base so calibration is not where it needs to be, including because there is a deliberate effort to fight against such calibration.
It still does seem way more fortunate than I expected a year ago.
It seems safe to conclude that to get the level of reaction we need from warning shots, at least under the current configuration of players, will require us waiting until after someone gets hurt. Potentially quite a few people.
In turn, then, the “good” news is that with this level of irresponsibility and failure (another way of saying ‘the base rate is incredibly high’) we have a good shot at getting warning shots where only a limited number of people get hurt, before quite a lot of people get hurt or everyone dies, if we can avoid hitting a point of no return.
So yeah, with this attitude, or without this attitude, the incidents are not going to stop, or stop getting larger and more frequent. But if we are not on track to solve the underlying problems, a high rate of warning shots is much better than the alternative.
Alert the Media
When the facts change, I change my opinion.
It could happen to you.