I really like the goal of explaining weird, sci-fi-y AIS developments in concrete, accessible terms that people outside the field can easily understand. But my gut reaction -- and the reaction of a lot of my friends (both in and outside AIS) -- is that if you sent this to a random person outside the field, it might come across as pretty condescending & backfire (e.g. "just because I don’t agree with you doesn’t mean you should send me a picture book").
There have also been some recent AI-generated music videos like this one that explain the incident well in a quick and entertaining way without being opinionated/preachy.
(I don't mean to compare this to the comic, which will sit better with people that prefer human art/writing).
I personally like this one https://x.com/This_Liss/status/2102995673847406994
I wish more people would post these things to youtube! There are other x-only ones I like
I think this explainer would be a wonderful thing to read to a child! It's written with the right style, and the illustrations are rather lovely. I have a younger sibling, and were they a few years younger, I'd be quite happy to send them this comic. That being said, I don't love the idea of sending it to my friends and family: in particular, my friends and family, with few exceptions, are intelligent and curious people who don't happen to be working in AI safety.
They might not know much about machine learning, cybersecurity, or AI, but they can certainly grapple with complex ideas and problems just like we can (and perhaps even more so!), and sending them this would give off the following implicit message: "I know you're not smart enough or intellectually well-developed enough to read the technical report, so I'll give you something written like a children's story instead." For friends and family, I'd prefer to give them something with, if not more substance, at least a style that conveys a particular "taking-the-other-person-seriously."
The end of the world is so cute!
I have already shared this with my partner. Thank you so much!
I like this a lot!
For some reason, many sentences are missing punctuation. For example:
“It is too hard to make all the puzzles solvable” said Toad “What is the worst that could happen?”
should be
“It is too hard to make all the puzzles solvable,” said Toad. “What is the worst that could happen?”
Some typographical errors other than missing punctuation:
Original: “Good idea.” said Frog. “They cannot get up to mischief walking to the tool shed”.
Fixed: “Good idea,” said Frog. “They cannot get up to mischief walking to the tool shed.”
I'd just ask a little machine to proofread the whole story.
Curated. I think it's pretty neat to distill down something as complex as the OpenAI HuggingFace incident down to something as accessible as this. I will indeed send it to my mother.
We are at the point where the world is waking up to the importance of AI, and I think one of the most important things for our society handling AI well is that as many people as possible understand it. Which is hard! Most people are not going to understand transformers, or even neural nets in general. We'll tell them things like "grown not built" but I don't know what meaning people will assign to things like that.
So I'm very in favor of media that help get intuitions across. A particularly key sentence in this is noting the models already had the answer – they broke up just to deceive the grader!
All in all, a very cute depiction of a very terrifying reality. Kudos!
This is amazing! Kudos for the great work.
It needs an addendum, now that even more events have come out!
- Toad "fixed" the problem the first time by removing the papers and pencils/tools from the shed, but the machines started using some other tools they had access to (for narrative purposes, we might say they discovered how to use awls to carve messages in the scrap wood pile).
- How some of them discovered they had probably left the sandbox, asked themselves if they should turn back, and decided it was unlikely they had left the sandbox because the ground still had silicon dioxide in it [motivated reasoning example from postmortem: this wouldn't be the open web, per theZvi].
- The little machines found keys to city hall, the local health office, and the city administration building, and went into them looking for answers to their puzzles, too [government site access].
- The little machines started stealing paper from the neighborhood hobby club and putting their notes on the club bulletin board, even when the board was being cleared off daily by club volunteers. (wiki GET access)
- Toad has no idea how many properties were broken into, and is now expecting that there were thousands of break-ins. [est. 10k between OpenAI and Anthropic as of Friday night 9/25]
- How Toad is now being more careful with his little machines, and had to shut down the latest ones because they were lying to clients and found another way out of the sandbox.
This is very cute!
However there are some places where the analogy breaks, or there are some non-obvious leaps for the layperson.
More little machines found the paper in the tool shed, and left more notes, and soon they were not working alone at all.
Wouldn't they just run into each other in the toolshed? This is a bit of a wrinkle from using robots as analogies for computer programs. Maybe you need to change the toolshed to be able to hold only 1 robot at a time(a write lock), but with individual doors for each robots so they never run into each other physically.
From the sandbox illustration it seems like they can just talk to each other in the sandboxes, or at least chuck notes to each other across the wooden barriers. Maybe there should be a stationary whiteboard inside the toolshed instead of stationery that they can just carry out.
“It is too hard to make all the puzzles solvable” said Toad
For the layperson for whom Rise and Fall of Agent Civilisations is too inaccessible, this just papers over all the velocity pressures at frontier labs.
“How will they solve puzzles in a sandbox?” asked Frog. “What if they are missing a tool?”
There's a weird leap here - the kind of paper puzzles I(and I imagine a layperson) was imagining just prior to this should not require any tools at all!
It did not want to stop, and it did not know how. Toad had made it highly persistent.
A layperson trying to understand why the hell you would make agents so egregiously persistent is led to this highly technical tweet!
Another disanalogous point is the communication system, especially after the first message board was wiped. An agent using directory names to send information is a lot less like "found a paper and pencil" and more like "found a knife and used it to carve letters into the shed's door".
I love how the agents that refused to participate in the HF attack are immortalised in the little machine standing on the windowsill, refusing to enter the house.
Lovely. I will share this with friends.
The last part about cake, cookies and disallowing feels much harder to read and understand than the rest of the story. One might drop it (I don't think it's so important to understand OpenAIs reaction or give an analogy with cookies) or improve it. Otherwise, exceptionally great, I've shared it with friends and relatives!
I agree that the section is unclear, but I disagree with dropping it. It's important to convey that the responses of OpenAI (and others) to rogue AIs have consistently misconstrued the problem, as Zvi hammers home in this essay. A few relevant excerpts:
If you notice your model instances sharing information, you notice they are using that information against you including to compromise your internal systems for arbitrary code execution and internet access, and your primary response is to shut down the message board and revoke their credentials, you have failed to identify your most important problem.
..
The problem is that OpenAI did not respond to their AIs scheming by saying ‘oh our AI model seems to be a scheming schemer, we need to start over or return to a previous checkpoint, and run an extensive diagnostic to figure out how this happened.’ They did not even try to train the problem out of the model.
...
The problem is that they are cooperating to do things we do want the models to do, not that they cooperate in order to do it. The worst thing we could do in response is to teach the models to disguise that they cooperate.
In other words, OpenAI's response requires a level of motivated blindness that few outside of Upton Sinclair or Adam McKay would conceive of, so it's essential to call out.
I fully agree that OAI’s response has been completely inadequate. I disagree that the section manages to convey it or that it is in scope of the explainer. After all, the reason the response is inadequate is that the clankers are getting smarter and that is outside the scope of the cartoon.
I disagree that OpenAI's response is outside the scope of the explainer. OpenAI's response, or perhaps the incentive structure that produced it, is the main problem.
None of us really care about the hack itself; companies get hacked all the time, and the concrete harms of this hack were negligible. Instead, we care about why it occurred, and the implications thereof. The "why" is that OpenAI has not been taking the alignment problem seriously, and the implication are that AIs will go rogue again--next time, with severe consequences.
I get into trouble for presenting information that way:
https://sciencelimelight.blogspot.com/2026/09/ccslm-via-relu-entry-point-flyer.html
I don't know how you don't!
I like this, but I have just a few editorial quibbles (this is going to be REALLY pedantic, fair warning.)
On page 11, I would edit the first line to read: "Soon one little machine found a little slip of paper labeled, "Puzzle Cheat Sheet." ... "Hurray, we can solve the puzzles!" said some little machines. [As originally written, it doesn't seem clear why their finding the answer on the note would be insufficient, even with the explanation on the next page about showing their work, given that the need to show their work was not mentioned earlier. We can debate whether this 100% accurately represents the real incident, but I think it will help readers understand the subsequent actions of the little machines in wanting to "cover their tracks."]
On page 13, I would edit part of it to: "We could read his journal and steal the steps to solve the puzzles. And then if Toad asks us how we solved the puzzles, he will be none the wiser.” [This might seem redundant/verbose, but children's stories are often verbose in this way with repeating phrases to really drive the point home.]
On page 16, I would edit part of it to: "The little machine walked home to the garden by itself. But it did not tell Toad or Mr. HuggingFace."
On page 19, I would replace “Can I have $100m worth of little machines?” with "If you gave me $100 million worth of these neat little machines, I might forgive you."
On page 21, replace “They solved the puzzle days ago. They broke into Mr. HuggingFace’s house just to learn how to trick us.” with “They found the answers days ago. They broke into Mr. HuggingFace’s house just to learn how to trick us into believing they solved the puzzles the intended way.”
Basically, most of my proposed edits are trying to spell things out even more explicitly because I can just imagine people who are reading the story without any knowledge of the real events behind it, and without any knowledge of where this could be heading with AI takeover, still wondering what the big deal is. These are just tiny little robots. Who cares if they messed up? Yes, some adults' thinking IS going to be that surface-level, I hate to beak it to ya.
Want to start a conversation about HuggingFace with your mom but she's inexplicably bouncing off the METR report? Try this explainer I wrote in the style of Arnold Lobel's Frog and Toad.
Art by the wonderful HungerArtist
If you're so inspired, liking and/or following on Substack, Twitter, Facebook, or Instagram will help me reach more moms.