So let me get this straight: A computer system carried out a sequence of actions that would be years-in-prison felonies if done by a human being. The system owners' response is "we're slowing down the speed with which we give this system new and more powerful capabilities, and talking with the victims." Am I missing something? I'm not so much worried about the "alignment" of the computer system; I'm worried about the alignment of the owners.
This isn't Terminator 2, folks; this is Tron.
The companies do have hundreds of billions of dollars that they could use to pay fines/lawsuits for quite a lot of these incidents, even if they were fully liable. And if they were fully liable, that would in some ways be the system working as intended. If they're producing so much economic value that they can afford to pay for all the damage they're doing, then maybe it does pencil on a society-wide level for them to keep going.
(Of course, at the moment the companies aren't clearly liable for this kind of thing! Which presumably contributes to them being more reckless. I'm just saying that, even if they were fully liable, they might not immediately do anything much more drastic than "we're slowing down the speed with which we give this system new and more powerful capabilities, and talking with the victims.")
But even if they were fully liable, this story wouldn't work for those existential/society-scale risks. The companies won't be able to internalize the harm of those, so I think we need further measures there. And this incident seems important as evidence and as a warning shot for those risks. (C.f. here.)
I agree that fines would not help the x-risk situation.
(For one thing, OpenAI already spends a whole lot of money paying off government people in exchange for unfair advantages and special treatment. "Pay the government more money in exchange for letting you keep doing business" is not an incentive structure, it's just more shakedown.)
I suspect that shutting down OpenAI would help, if it's possible.
Oh, I think fines and liability might help the x-risk situation on the margin via incentivizing more safety work.
"Pay the government [or people suing you] more money in exchange for letting you keep doing business" seems like an incentive structure to me if the money paid is proportional to how much harm you're doing. Which isn't going to be doable up to x-risk level, but it could be doable below that, and there's some overlap in the work you want to be doing for both.
The main point of my first comment was just that I think that risks of incidents at this scale isn't a good reason for OpenAI to stop. I think the importance of this incident is that it's evidence about misalignment risks that could do much more harm in the future as the AIs get much more powerful, if they stay misaligned. (It's possible you agree with this and I just misunderstood your first comment.)
I think pushing for regulation treating AI deployers as responsible for actions taken by their AI as if they had intentionally done the same action themselves could be hugely helpful at slowing down AI deployment until they are reasonably sure it's actually safe.
Remember the debate on whether or not we would be able to see significant 'warning shots' before doom? And how they would give us an opportunity to steer in another direction before it's too late?
Well, if this one doesn't count as a warning shot then what even does?
To get a sense of what the NYT's audience (mostly mainstream liberal) thought about this story, I had Claude do a mini-thematic analysis of the comments whose content is accessible to me (which appear to be ~181 comments all of which have a minimum of 2 upvotes). Here were what Claude Fable thought were the commonest themes expressed, along with the number of comments and aggregate upvotes of comments expressing those themes (each comment assigned to only one theme), sorted by total upvotes:
Claude Fable's thematic analysis
Theme | Total comments | Total upvotes |
Existential alarm: an AI escaping human control portends catastrophe, possibly for civilisation itself | 17 | 875 |
OpenAI / the industry is reckless, careless, incompetent or unethical, and doesn't understand its own systems | 12 | 620 |
Urgent government / international regulation is needed (guardrails, privacy laws, non-proliferation-style regimes) | 12 | 534 |
Doom expressed via sci-fi allusion / gallows humour (HAL, Skynet, Terminator, WarGames, Frankenstein, Colossus) | 25 | 486 |
The story is under-covered and deserves front-page prominence | 5 | 389 |
Scepticism of the official account: a publicity stunt, pre-IPO hype, or a deliberate attack on a rival | 18 | 262 |
Wider lesson: all connected infrastructure (banks, vehicles, medicine, defence, records) is now vulnerable | 10 | 256 |
The containment/test design was the failure: it should have been air-gapped; "that's not a sandbox"; "just unplug it" | 11 | 232 |
Fury at the phrase "a cost of research velocity" as revealing OpenAI's warped priorities | 7 | 159 |
Pure jokes and levity without a strong analytic position | 8 | 144 |
Cynical business-model critique: create the disease, then sell (or get bailed out for) the cure | 5 | 144 |
Legal accountability: unauthorised intrusion is a crime; humans/companies must bear liability and compensate | 5 | 114 |
Miscellaneous one-off views and thread side-chatter (IPO disclosure, gender critique, political asides, thanks) | 9 | 104 |
Biosafety analogy: AI "gain-of-function" research / a lab-leak-style escape (Covid parallels) | 3 | 81 |
General distrust/contempt for the tech industry and its executives (incl. from self-described tech veterans) | 8 | 75 |
Technical questions and explanations about how this could happen (incl. the NYT reporter's clarification) | 4 | 75 |
Development should be shut down, paused, or drastically slowed (moratorium) | 5 | 74 |
Anti-anthropomorphism: "going rogue" is misleading hype; it's software doing what humans built and prompted | 5 | 41 |
Pushback on the stunt theory: this is damaging press, not marketing, and the theory doesn't add up | 6 | 39 |
Noting the irony that Hugging Face relied on a Chinese open-weight model for its incident response | 2 | 38 |
Political cynicism: the current US government will not act | 3 | 32 |
Measured pragmatism: AI is here to stay; adapt (e.g. study cybersecurity) and manage the chaos | 1 | 3 |
Second group of headlines on reuters.com, after Iran / Houthis: OpenAI AI models went rogue during testing, triggering 'unprecedented' breach at startup
Holy crap. If AI leaders don't learn from this we're all doomed.
(Also why are these evals not running on an airgapped system? That would make sense if you're taking off the safety guardrails but of course a sufficiently powerful AI may be able to exfiltrate anyway.)
They aren't airgapped because air gapping is expensive (mostly in lost productivity) and annoying, and insufficient in and of itself (and probably not the most pressing security failing),
and the companies / the employees don't wanna, they wanna race.
As far as I understand the system was not air-gapped in order to allow the models to install packages. They then used an escalation of privilege to gain control of the cache proxy and therefore also gain internet access.
openai.com/index/hugging-face-model-evaluation-security-incident/
What do they mean by "contained the models". Like - there was no weights transferred, presumably, nor I hope would the attacking models have access to those. Presumably this attack was via a bunch of API calls and new code transferred to Huggingface's system? So what did Huggingface do to "contain" something fully, that they don't have access to?
Thanks!
This does need a link to the OpenAI blog post. Here is the link: https://openai.com/index/hugging-face-model-evaluation-security-incident/
(This might be another instance of LessWrong linkpost functionality being unreliable lately. It did fail for me in my last post a week ago.)
Link: https://openai.com/index/hugging-face-model-evaluation-security-incident/
From the OpenAI blog post:
(emphasis added.)
Yesterday, OpenAI disclosed that some of their internally models were misaligned. Today, they disclosed that "a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model" had compromised HuggingFace infrastructure in the course of running some OpenAI internal cyber evaluations on ExploitGym.
These cyber evaluations were supposed to be run in sandboxed environments, with internet access limited to installing packages, then:
The exact scope of the incident is unknown