I've commented on this announcement on twitter and in the Constellation slack (from which I've resigned my membership, as per below). To quickly summarize: the AI safety community is structurally blocked from reliably improving the world unless it becomes much better at handling adversarial dynamics. In particular, the most important lesson we can learn from the last decade is how and why "AI safety" was so heavily captured by AGI companies, and what it would have taken to prevent that. The most basic building block is the willingness to call a spade a spade, and honestly discuss what's going on with AGI companies; I interpret Paul as committing via his association with OpenAI not to do so. Of course Paul is only one part of Constellation, but his worldview has been (and I expect will continue to be) the single biggest influence on the broader strategy that Constellation is pursuing, which suggests that most reasonable thing to do is to disaffiliate.
The text of my tweet:
I am sad and disappointed to hear that Paul is joining OpenAI's board. Being affiliated with OpenAI has historically led AI safety researchers (including both Paul and myself) to behave in low-integrity ways. I pers... (read more)
[Paul is] the closest thing there is to a leader of what I’ll call the “pragmatic AI safety” cluster
Is that true? I haven't been in Constellation since a long time, but my impression was that Paul largely stopped commenting on LessWrong and Constellation Slack around the time he joined the US Government. Is he still a particularly influential voice in AI safety? I thought he traded away his voice for USG involvement years ago, and at that point he might as well get involved with OpenAI too. In general, it doesn't seem that bad policy to me that some thoughtful people should try to specialize in joining powerful institutions and talk sense to them at the cost of their public voice, while others should remain fully independent and try to become honest and unbiased public thought-leaders.
In general, it doesn't seem that bad policy to me that some thoughtful people should try to specialize in joining powerful institutions and talk sense to them at the cost of their public voice, while others should remain fully independent and try to become honest and unbiased public thought-leaders.
Isn't it crazy that our world makes these choices mutually exclusive, and on a meta level, everyone just takes it in stride?
BTW who are some remaining public thought-leaders with views similar to Paul's? I often wonder "what does Paul think about this development or idea?" and I'm not sure whose posts/comments to look up for the closest substitute. I guess Geoffrey Irving comes to mind but he doesn't post/comment nearly as much as Paul did.
Here's a list made by Perplexity[1], but none really fit a combo of prolific poster/commenter, independent voice, and focus on scalable pragmatic alignment, that Paul represented before joining USG:
Some notes on public communication, which I expect the LW audience is especially interested in:
If anyone who thinks this comment is somehow inappropriate (or doesn't get why it's amusing) would like to explain why, I would welcome that.
One of the top 3 things that has made Paul's career has been writing a lot in public about this subject in a way that is respectable, never rhetorically alarming, and able to serve as an optimistic counterweight those sounding the alarm. So it is funny in the immediate weeks after the NYT has frontpage pieces about rogue AIs committing crimes and Bernie Sanders announcing a bill to ban superintelligence, to see his one tonal note be that he has softened his writing even further.
I strong downvoted the parent as it seemed unproductive and distracting. I think it's important context that Paul is making the tone softer than what it otherwise would have been given he is a director, and it would be useful for people to extrapolate to what the post otherwise would have been.
able to serve as an optimistic counterweight those sounding the alarm
FWIW, I don't think Paul served a large role as an optimistic counterweight to those sounding the alarm, at least in public.
So it is funny in the immediate weeks after the NYT has frontpage pieces about rogue AIs committing crimes and Bernie Sanders announcing a bill to ban superintelligence, to see his one tonal note be that he has softened his writing even further.
I think there are communication compromises you'd typically take when joining the board of a company. I think there is an open question about the right strategy (both tactically and given typical norms) in circumstances like these.
I read this post as relatively stark and concerned compared to Paul's earlier communication. So while the effect of being a director may have softened his writing relative to what it otherwise would have been, other effects (e.g., the immediacy of risk) may be making it significantly more harsh.
FWIW, I don't think Paul served a large role as an optimistic counterweight to those sounding the alarm, at least in public.
Which time period are you thinking of?
I definitely never perceived Paul as an e/acc, if that's what you're imagining as an optimistic counterweight.[1] But I've been following this stuff since ~2010, and working in the Bay on it since 2016, and my experience of much of that time was there as being two main thought leaders, Eliezer and Paul, with Eliezer characterizing the pessimistic view that thought alignment was difficult enough that you needed agent foundations to have a good scaling story, and Paul championing the view of relative alignment ease and thought we had a reasonable shot of aligning prosaic AI.
And like, if you read that post (originally from 2016, published on LW in 2018), he definitely doesn't seem happy about the situation:
... (read more)I feel pessimistic about human prospects in such a world.
So I am extremely unhappy with “a significant risk of trouble.”
We might hope that this situation will change automatically as we build more sophisticated AI systems. But I don’t think that’s necessarily the case.
Like I said, I think this would probably be OK, but it o
Paul probably did serve as a substantial public counterweight to Yudkowsky-style pessimism within the AI-safety community
This seems right but isn't how I interpreted Ben's claim; I interpreted Ben as saying that Paul was a "optimistic counterweight" in the broader public discourse about AI safety (which is who one would naturally "sound the alarm" to), which is obviously different than the question of how he fit into the discussion within the AI safety community!
I slightly dislike the comment just because it’s ambiguous whether you‘re playfully ribbing Paul for his idiosyncrasies, or being derisive because you’re in all seriousness unhappy about his comms.
At first I thought it was meant to be the former but accidentally came off as the latter, but after reading the thread now I think it’s meant to be the latter?
Either version would be a fine comment imo but the ambiguity feels slightly annoying (perhaps because people sometimes deliberately make such comments ambiguous as a rhetorical tactic? not saying I think that’s what you did though)
(fwiw, I think the things that make biting-culture work among friends is that they're... like, on the same page about what page they're on, so when they make snarky remarks about each other it's clear what kind/flavor o snark it is. I don't think it translates across people/groups that you know less well and it's more ambiguous whether you think each other are like slightly fucking up or doing really egregious things)
I agree that talking about tribes is generally bad, and it would be good if we could go back to doing less of it.
But I think that in the recent past, John Wentworth, Richard Ngo, Eliezer and others have started to ignore this norm, railing against the moral failings of a nebulous outgroup, "the EAs", and getting wildly upvoted for that with minimal pushback even when their comments are pretty low quality imo.
(My least favorite example is this comment from John, calling a nebulous outgroup immoral and underperforming a five year old, while not providing any specific examples of people lying, or how to navigate the tradeoff between The Obvious Right Thing To Do "spread true important things, don't strategically hide your views", and between the view that it's e.g. bad for Thomas Kwa to go to OpenAI to investigate RSI and publish about it. I think the entire comment is just a naked status attack on the other tribe, with very little justification provided, but it sill got highly upvoted.)
Maybe there is a good reason for John, Richard, Eliezer, etc to start loudly blaming the nebulous other tribe: sometimes coalitions break down and it's good to name things as they are.
But I think at t... (read more)
Recap for those not tracking the details: my understanding is that the nonprofit (“OpenAI Foundation”) owns 26% of the shares in the public-benefit corporation (“OpenAI Group PBC”). However, the non-profit board has special governance rights that give it complete control over who sits on the PBC board: it appoints all of its directors and can replace them at any time.
Confusingly, the two boards have so far been nearly identical, with 9 people sitting on both, including Bret Taylor as Chair of both boards, and Sam Altman also serving on both (being the only PBC executive to do so). Zico Kolter (Professor at CMU and Head of its Machine Learning Department) was the one exception, serving only on the nonprofit board (and as a non-voting "observer" on the PBC board). Paul Christiano now appears to be the second person appointed only to the nonprofit board.
I guess if there are subcommittees, perhaps more people will be announced soon. We’ll see!
This seems like a position that should not be accepted lightly. Do you have reason to believe the SSC and the board will in practice be a meaningful check on OpenAI?
A twit by the founder of an AI safety advocacy organization states that
the SSC in particular oversees model launches and can delay releases for safety reasons
I have never heard this before and is outside the usual remit of boards. This would be a strong positive update for me about OpenAI's governance having any teeth (though it will not prevent an extinction-level threat, because of course most of the risk comes from training the model and internal runs, not releasing it to the public). Anyone got any more information on this?
and keep in mind that the company can just ignore its governance mechanisms
I am kind of confused about the linguistic construction of "can delay releases for safety reasons" and "the company can just ignore its governance mechanisms". It is not currently clear to me whether the board can actually delay releases for safety reasons (though it might be in some sense supposed to be vested with the power to do so).
Yes I believe they have the relevant formal power. At least it's in the MOU between the California AG and the OpenAI legal counsel.
The SSC has and will continue to have the authority to require mitigation measures—up to and including halting the release of models or AI systems—even, for the avoidance of doubt, where the applicable risk thresholds would otherwise permit release. The NFP will provide advance notice to the Attorney General of any material changes to the SSC’s authority.
[SECOND EDIT] Ben raised the excellent point that Paul had presumably signed OpenAI's non-disparagement clause in 2021 and had never mentioned it publicly, which is a concern when it comes to predicting future whistleblowing willingness. Paul replies at length below, and I find his reply to be pretty reasonable.
While I believe that it is unethical to work for OpenAI in any technical capacity (it will be twisted into capabilities progress), I think the role you are taking is defensible, and I trust you in particular to leave (and publicly state why) if you discern otherwise.
Good point; that's evidence against Paul's willingness to blow the whistle when called for, which is (as they say) load-bearing here.
@paulfchristiano, have you ever publicly addressed the topic of whether you signed a non-disparagement agreement when you left? Opus 5 didn't find an example.
I haven’t addressed it publicly. Here’s a timeline and some root cause analysis:
Thank you for the detailed explanation. According to this summary, the standard agreement with OpenAI went beyond non-disparagement and included a secret life-long agreement, "not to interfere with OpenAI’s relationship with current or prospective employees, current or previous founders, portfolio companies, suppliers, vendors or investors."
If the signee publicly admitted to making this agreement, then (per other agreement portions) OpenAI could remove all monetary value from the signee's OpenAI shares.
Was this restriction or similar also part of the agreement you signed? Did you hold OpenAI shares when employed at NIST? What about currently?
I did not have any OpenAI equity at the time I joined NIST and do not have any now. I signed the standard agreement you referenced, though I don't agree with that post's legal or practical analysis.
Can you describe what you think the organization is doing by bringing you on? What has not been done, that will get done now that you're in the room?
What exactly is the Safety and Security Committee doing and how much power does it have? Can the board employ safety researchers that work for the nonprofit and that aren't employees of the for profit?
To what extent does this help us with a global development pause?
I am excited to be joining the OpenAI nonprofit board, serving on the Safety and Security Committee to support safety oversight.
Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.[1] I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level. I’m joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk.
The SSC has an important and challenging role in overseeing risk management at OpenAI, and I hope to help provide expertise and assistance in a critical moment. My joining is not an endorsement or criticism of OpenAI’s safety practices in particular; I hope that all frontier companies strengthen safety oversight and I am excited to work on this at OpenAI. I believe that the rest of the world should judge OpenAI, and all AI developers, by externally verifiable behavior and results.
In the rest of this post, I'll explain why I believe loss-of-control risk is now acute and how I think about the current situation in the AI industry.
First, automated AI R&D could lead to a very rapid acceleration in AI capabilities very soon. OpenAI has predicted that we might have capabilities sufficient to fully automate AI research within 18 months; my personal forecast is extremely uncertain and I think it could easily take anywhere from several months to several years.
Full automation of AI R&D means that improvements in training and algorithms can directly increase the quality and quantity of automated AI researchers available to do additional research. Existing evidence is very uncertain but suggests that this positive feedback loop might be strong enough to overcome diminishing returns and compute bottlenecks, leading to a rapid intelligence explosion. If this happens, then within six months of full AI R&D automation we could see more algorithmic progress than has occurred since the development of the Transformer nearly a decade ago. I believe this would result in superintelligent AI systems.
Second, we currently train our AI agents with RL to get as much reward as they can. It has long seemed theoretically possible that this could motivate AI agents to undermine human control, seek power and resources, and cover up their tracks in pursuit of misaligned goals correlated with reward. Public evidence from recent incidents suggests that this is not just a theoretical possibility.
An intelligence explosion would greatly exacerbate risks from misalignment, both by making the technical problem of alignment even more difficult and by rapidly raising the stakes for failure. Many researchers and leaders at OpenAI and across the industry have expressed concern that rapid recursive self-improvement is not consistent with safe development; I resonated with this recent post by OpenAI's chief scientist Jakub Pachocki on this topic.
If we build superintelligence without more robust alignment I expect we will permanently lose control of it. If that happens then most people could die. I believe we would need domestic and international coordination to ensure global consistency and reduce risk to an acceptable level.
That said, frontier AI developers have a lot of power to unilaterally improve the situation and lay the groundwork for stronger coordination. Developers can improve safety mitigations (including slowing development as necessary), transparently share evidence about risk and the effectiveness of their mitigations, and work towards shared safety standards.
I am encouraged by other members of the SSC, as well as the rest of the board and leadership, taking these issues seriously. I look forward to working with them to help OpenAI raise the bar for its safety practices.
If I had to quantify my uncertainty I would estimate an all-things-considered risk of 4% over the next year and 15% over the next three years. These numbers are a way of stating my subjective beliefs and communicating roughly how big I think the problem is, not a claim to have a model that produces precise or stable estimates.