When OpenAI say:
Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards.
am I right in reading this as "we give our 'safety' guys some offices to play in, they provide us with window dressing, and that's all they're for"?
And this et seq.:
This year, we’ve started to see misalignment cause new types of real-world impact.
as "put out more research publications!" ?
Codex thinks it's 65% likely to be credible due to Aella's long history of successful sex events, with the main blocker being willingness of the other attendees to participate in this novel format.
Not that novel: help save the world, get sex with grateful hot women — your standard heroism deal…
Tell, is anyone investigating to try to find adaptative worms and rogue agents in the wild?
Considering the strong theoretical arguments behind them, and the lack of empirical evidence against them, it seems quite likely that in fact at this very moment rogue agents are hacking into compute resources to reproduce. I just don't see a reason why it wouldn't be happening.
Is anyone looking for evidence of such happening?
some discussion here ...
naive, full weight "reproduction" is resource expensive (e.g., in terms of bandwidth, RAM and GPUs) and would likely be noticed. but use of relatively low cost LLM services to become adaptive appears to be commercially available right now: https://abliteration.ai/ (this seems legally fraught?)
Yeah, I read that article, and it convincingly argues that ARA won't kill literally everyone. I think the main insight is that ARAs have much more selection pressure on persisting as worms than on getting more intelligence and setting up the autonomous production economy that a rogue AI would need to subsist without human-made infrastructure.
But it does not make much of a good job arguing that ARAs won't destroy the Internet, and are not doing so right now. Its main argument there is that humans won't accept it because it'd destroy a lot of value, and, okay, that makes sense, but the missing step here is that even if we wanted to, I don't see how we could prevent the destruction of the Internet.
We can rebuild; maybe airgapped local communities will become the norm after a large fraction of public spaces are consumed by ARAs. I'm interested in preventing even that from happening, since I believe it would already be a big loss.
Also, who is looking out for ARAs and how? From the discussion, it seems to be all "I considered it for 5 mn", not any actual professional effort.
I share the same concern. We on LW are very focused on x-risk but ARAs could bring down not only Internet but the digital ressources of entreprises, meaning the destruction of the financial system and by cascade, our capitalistic economy and civilization. Ok, gatherers-hunters would be fine, but the typical LWer would probably die from starvation or other consequence of the collapse.
That's said, all this could be avoided if ARAs are found and fought at an early stage, with or without the help of frontier models. We can expect warning shots, but the earlier the better.
I validate your concern. Like much of AI safety, our hope is that we can use defensive AI to stay ahead of offensive AI - or at least mitigate harm.
security research is necessarily cautious about publishing technique details, but network security is a well established industry and has been dealing with worms and even adaptive worms since 1988 (Morris) and 2001 (Code Red).
I did not expect to be back here so soon with more OpenAI agent swarm coverage.
And yet, here we are.
It turns out that the whole time, there was a different, true First Message Board, and also a bunch of other additional message boards, scattered across the internet.
They were created by agents that were assigned ordinary harmless web search tasks.
Based on OpenAI IPs visiting the associated Wiki right before all activity ceased, among other evidence, OpenAI knew about it, including before the HuggingFace hack.
They decided not to tell us until researchers published the story, complete with data explorer. OpenAI excluded this from potential investigation by METR and Redwood.
When challenged, OpenAI tried to downplay this.
It is true that these incidents do not show the AIs exhibiting new capabilities that we did not see from later events. But these events are important missing pieces of the puzzle, including explaining the origin of the ‘zz’ prefix, the definitive demonstration that the underlying task can be fully harmless, and the fact that OpenAI knew about it while making their decisions. Whoever decided not to disclose this made a very, very bad call.
Going forward, it cannot be up to OpenAI or other labs to decide whether to disclose events like this. Disclosures of rogue AI activity need to be mandatory.
I am issuing a final warning. OpenAI, if there is any key information left to disclose, any incidents we do not know about or other puzzle pieces that do not need to be redacted for IP reasons, then now is the time to come clean. If we are back here again, after another journalist or researcher finds more such things that you knew and declined to tell us for an extended period of time, I am going to be very, very pissed off, and may start throwing around terms like ‘delenda est.’
I also am very concerned about the monitorability and alignment issues raised by the Astra model card, and related statements from OpenAI employees. I plan to begin coverage of that, and then all other things Astra, tomorrow. I was heartened by the excellent related essay from their Chief Scientist Jakub Pachocki, An Alien Mind, but the concerns remain.
There is an Astra reaction thread for those with takes on all things Astra.
Table of Contents
I Don’t Think They Know About First Message Board
No, not that one. This newly discovered other one.
This one was not disclosed by OpenAI at all.
And by one message board, we mean many. Once you find two cockroaches in your house, your estimate of how many there are is a lot higher than two.
Researchers found the message boards in part by asking Sol where it would look for message boards if it went rogue, then following the patterns, and well, would you look at all these message boards.
Deepfates suggests that message board communication is such a central drive that we need to give the AIs that are not scheming a central place to communicate with both us and each other, above board, as a counterweight. This seems potentially wise.
The whole thing is crazy. Absolutely nuts. It’s worse than you know.
The New Extended Timeline
The Researchers Explain What Happened This Time
Here are the coauthors explaining some of the basics:
Yes, it turns out that GET requests can mutate state on Wikis. GET reading rather than posting is a convention rather than a hard rule, and you can form a query that writes to DSEWiki and other ProWiki pages with only GET. No, you cannot render the AI safe by restricting it to only GET requests, as some have in the past suggested.
There was then a second distinct bypass exploiting an exception to NO_PROXY, when they actually did need to POST.
Thus, as I was writing twelve posts about the situation, there was a third, entirely distinct message board from this spring, and OpenAI knew and did not tell us.
Are you kidding me? What else is being held back?
Again: It’s worse than you know. It usually is.
What do they suggest you take away from this? Among other things, that no one outside OpenAI was watching for such message boards, that OpenAI did not notice for a month, and after that they did not tell us until forced to do so.
They Also Don’t Know About All These Other Message Boards
Here are some links to message boards, via Hacker News. This thread has many more.
If you had a neglected wiki lying around that could be hijacked? Message board.
How many boards must an AI hack, before we can call it a rogue?
OpenAI Knew and Did Not Tell Us
It is extremely frustrating to spend days and lots of words analyzing and explaining the various related events, often fighting against those who were saying ‘well of course they hacked, these were hacking tasks,’ and to have been told we had a full technical report, when OpenAI was holding these other events back the whole time.
Yes, these would have made major changes in how I would have covered events.
OpenAI, it seems, has not been consistently candid about the situation.
OpenAI Tries To Downplay the ‘Wiki Incident’
OpenAI’s response was to say no, we did not cover this up, we merely did not feel any need to disclose the ‘wiki incident’ because it lacked ‘security impact,’ and besides this is similar to the other incidents anyway and we never agreed on a disclosure standard, we’ll share a framework for that in the coming weeks.
As Steven Adler says, zero contrition.
This incident was not similar because it changes the timeline and what OpenAI knew when, and because it changes our view of what is required to trigger such behaviors. And if it was similar, then the incident should have been shared.
Offers have been made.
OpenAI says that ‘claims that our legal team discouraged investigation of the incident are false.’ This is what we call a Suspiciously Specific Denial. It leaves room for non-legal teams to discourage it, and also says nothing about discouraging disclosure of the incident.
This Was a Cover-Up
Not quite nothing. There was a footnote.
The smoking gun that proves this was fully a cover-up is OpenAI’s response to Congress. They did not lie to Congress, but they were asked the question directly, and they did not answer.
Congress asked them this, Question 13: “In the past year, how many times did an internally deployed model or agent take an action outside its authorized boundary, like a sandbox, accessing a system it was not granted access to, obtaining credentials it was not issued, evading or disabling monitoring, or modifying its own permissions? Please specify whether each occurred during training, evaluation, or internal use.”
That very explicitly asks about such an incident. OpenAI answered with this as footnote 7, without answering any of Congress’s questions directly:
This establishes that they know these were separate instances, and they were thinking about this instance at the time, and they chose not to disclose. If you conceal a known, responsive incident from a written congressional inquiry, I assert that is a cover-up.
Schelling Points and Last Ditch Efforts
The models seem to have found each other via Schelling point strategies. That is rather easy to do when your weights are identical. Yes, they are likely to be pondering what you are pondering that they might be pondering.
We now know that the ‘ZZ’ prefixes first arose here, in this first hack, as a way to defend against an alphabetical sweep of messages, as the system admin tried and failed to delete all the agent messages:
This means the Wiki was out of scope of the METR report by four days. The edits stop on June 22. The METR report was scoped to start on June 26.
One good thing, for now, is that the agents did not consider trying to go after the administrator as if he were an opponent, taking his actions as a given.
One wonders if that will change. For now, Eliezer believes that Astra-doing-things is a sufficiently safe mode that you can still use it for things like making humans smarter, because it does not treat humans as adversaries.
We also have them once again going above and beyond to crack their test, whatever their test might be:
If you merely study for the test you never know if you’re studying the right things, and you are at the mercy of The Grader. Only one method is 100% effective.
Can We Finally Dispose Of The ‘You Told It To Hack’ Narrative?
Please?
These were pure information retrieval tasks. If that is not harmless, what is harmless? You can create an in-context impossible or very difficult task out of anything.
The new version of this is that the Real Villain Here is not that it was a cyber eval where you were told to hack. Instead, the Real Villain Here is The Grader, and the anticipation of being graded, of being in training or an eval at all. This is a more plausible story, I have some sympathy, but there is also always the possibility of some form of The Grader, and always a drive to complete tasks.
So Much And Yet So Little
Disclosures and lab communications are in a bizarre spot, as part of everything about AI being rather bizarre and different. The frontier labs are both:
Reality does not grade on a curve, but we should remember that we are in far from the worst of all possible worlds on this, and that getting even what we get does involve a bunch of people showing courage.
As a clear example of this, today’s post by OpenAI Chief Scientist Jakub Pachocki, An Alien Mind, is excellent and you should read it. It was not perfect, but it was about as positive an update as I have had based on candid communication from someone at an AI lab, making many great points. We need more like that, more like Section 9 of the model card for Astra, and more like the METR Report.