I feel like this is giving way too little weight/salience to humans being really bad at strategy and philosophy in general, and in particular MIRI being bad at strategy and philosophy. You talk about Holden being misled about AGI for 8 years, but don't mention MIRI planning to build recursively improving Friendly AI with a small team and potentially just 1 philosopher, for a comparable amount of time. If OpenPhil had funded MIRI more in 2016, it would have been funding them to attempt this!
Just in course of searching for my name in the comments section of Holden's Thoughts on the Singularity Institute (SI), I came across more examples (of humans being bad at strategy and philosophy):
EDIT: rewritten to try to think through things a bit more clearly.
Returning to this comment, I notice myself being quite frustrated by it being the most-upvoted and first-appearing comment, given how little it engages with the actual content of my post. I do think Wei's point about MIRI is getting at something important, and I appreciate the evidence he gives in a later reply. But I feel some sense that.... hmm, maybe that I have to do a bunch of somewhat adversarial cognition to figure out how much to believe him? Like, Wei is so focused on demonstrating that humans in general (and MIRI in particular) are bad at strategy and philosophy, that I expect he's probably overstating the extent to which MIRI's strategy is actually well-summarized as trying to implement RSI themselves (which the disagree-voters on the MIRI claim also seem to think).
And even if I believed him on that point, I don't know how he thinks that should affect any of the specific points I made in my post—he's leaving all of that interpretative labor up to his readers, in favor of spelling out more examples of MIRI's mistakes which may or may not be relevant to my post. Because of that, the comment feels like a cha...
MIRI's attempt to solve technical alignment didn't make the global situation significantly worse whereas OpenPhil's funding of OpenAI and other "safety" work did according to some writers I admire.
Do you really think we'd be better off if no competent team with funding had made a sustained attempt to solve technical alignment? Alternatively, do you believe that there have been other competently-led sustained attempts to solve it outside of MIRI?
In an interview published on Youtube, Eliezer said, "I did my crying in 2015", by which he meant that that is when (presumably after the founding of OpenAI) he realized the global situation is hopeless (because superintelligence would arrive before a solution to alignment) which makes me wonder how you arrive at your belief that "if OpenPhil had funded MIRI more in 2016, it would have been funding them to attempt" "to build recursively improving Friendly AI with a small team".
You talk about Holden being misled about AGI for 8 years, but don't mention MIRI planning to build recursively improving Friendly AI with a small team and potentially just 1 philosopher, for a comparable amount of time.
I'm mostly not blaming Holden for the 8-year thing, I'm blaming him for how immediately after the 8-year thing he turned around and made a very similar mistake.
If OpenPhil had funded MIRI more in 2016, it would have been funding them to attempt this!
Can you provide a source for the claim that they were still planning this in 2016?
Either way, though, it seems really crucial that MIRI was doing the kind of plan that failed gracefully. That is, the first step of this plan was solving agent foundations (rather than throwing a lot of money and compute at stuff they didn't understand). And conditional on solving agent foundations, the strategic landscape would look pretty different (and they would have become much less strategically confused because they would have a better understanding of agency).
How does the idea of "strategic reasoning that's wrong but fails gracefully" fit into your ontology? Because to me the "failing gracefully" thing is actually the biggest single ...
Can you provide a source for the claim that they were still planning this in 2016?
It's based on the earlier stated plan (around 2013), no clear retraction of that plan (that I can remember or find) until much later, and @jessicata (who worked at MIRI 2015-2017) telling me some time after 2016 that MIRI was trying to recruit people for a serious push towards AGI (I think she was trying to get me to speak up against it at the time).
Perhaps worth clarifying that by 2016 MIRI was in the process of (or perhaps largely completed) moving away from CEV-style Friendly AI towards "Task-directed AGI" meant to perform a "pivotal act". (But I think this doesn't affect my point much.)
Either way, though, it seems really crucial that MIRI was doing the kind of plan that failed gracefully. That is, the first step of this plan was solving agent foundations (rather than throwing a lot of money and compute at stuff they didn't understand). And conditional on solving agent foundations, the strategic landscape would look pretty different (and they would have become much less strategically confused because they would have a better understanding of agency).
Suppose they did solve agent foundations and that...
An analogy here is expected utility theory, which totally changed how many people think about decision-making.
Most people don't think that solving agent foundations would lead to a new theory anywhere near as powerful as expected utility theory, but that's why I cited that leading agent foundations researchers would agree, because they're the ones who can see most clearly what a path to such a theory would look like.
I think Roko posting the thing was ok. It was in the same class of things being discussed at the time, like Rolf Nelson's AI deterrence and so on. Eliezer overreacted and caused a Streisand effect, without it only a few of us would even remember it today.
To add color to the point about me: I was a research associate at MIRI then (then called SI). I'd joined in the hope of doing decision theory math, but found that there wasn't as much math I liked happening inside. However, there were many email discussions about saving the world, which I first tried hard to follow, but then they became just really overwhelming for me. That's the background of my remark. Later that year I left the program. Maybe you're right and this all was a failure of strategy on my part :-)
I think Roko posting the thing was ok. It was in the same class of things being discussed at the time, like Rolf Nelson's AI deterrence and so on. Eliezer overreacted and caused a Streisand effect, without it only a few of us would even remember it today.
Well, my opinion is the opposite.
Trying to invent an infohazard and then publishing it... is stupid if you believe it, and an asshole move even if you don't. And perhaps this was destined to happen sooner or later, but as the article says, just because a slippery slope to hell exists, you don't need to jump on it enthusiastically.
"You should never delete a controversial comment on your own blog" seems like a norm that did not exist before the basilisk, and was invented afterwards as a rationalization. "Streisand effect" refers to an attempt to remove content from other spaces -- it would be an appropriate term if e.g. Roko published his thought experiments on Twitter, and then Eliezer tried to get his Twitter account banned.
Moderating your own blog was -- and still is -- a perfectly normal thing to do. For example, Scott Alexander deleted dozens of comments on his blog, and nothing remotely similar happened. What Eliezer overdid wa...
"You should never delete a controversial comment on your own blog" seems like a norm that did not exist before the basilisk, and was invented afterwards as a rationalization. "Streisand effect" refers to an attempt to remove content from other spaces -- it would be an appropriate term if e.g. Roko published his thought experiments on Twitter, and then Eliezer tried to get his Twitter account banned.
To correct the record, LW wasn't Eliezer's personal blog at that point. It was a community blogging site / forum, and Roko's post was a top-level post, not a comment. The Wikipedia page on Streisand Effect cites "Twitter CEO Elon Musk banned the Twitter account @elonjet" as an example, which seems pretty comparable to Eliezer's action.
(I deleted the first version of this comment, in order to confirm my memory that it was a top level post, not a comment on one of Eliezer's posts, but yeah, it's documented in the LW wiki.)
Yeah. Maybe it wasn't even due to that specific topic, could've been anything else, like knitting. There was just a lot of emails about it (several every day for months?) and for some reason it felt really hard to follow for me, on top of my work at Google at the time. So then it flipped around to "don't wanna talk about it, don't wanna meta-talk about it, just make it go away". I'm sorry the backstory isn't more dignified.
My concern is that the alignment community had a plan to make good outcomes more likely (differentially advancing alignment over capabilities) but has mostly pushed the world in the opposite direction, while some parts of it gained a lot of power by doing so.
Not sure how you're assessing the counterfactual. It's easy to say "here, alignment research accelerated capabilities!" but what does the ratio look like if there had been no public attempt at doing alignment research? (Or if ratio is not the right thing, how else to assess?)
I also think the strategic focus on "capabilities bad!" is in tension with the praise for scientific and conceptual progress:
More generally, one touchstone I’ll be referring to throughout this sequence is the idea that scientific progress proceeds by developing insightful new concepts
Unfortunately, most alignment research is no longer even aiming towards the kind of scientific progress I describe above.
So the kind of conceptual thinking that I’m praising in the rationalist community is what I’d call the generative part of science
Since... of course, if this were true, then alignment research (including agent foundations) would accelerate capabil...
I also think the strategic focus on "capabilities bad!" is in tension with the praise for scientific and conceptual progress
As Adria also points out below, I'm primarily using this as an example of alignment leaders pessimizing their stated goals, so that we can then create common knowledge around something being wrong with the way they make strategic choices. (Note that "pessimizing" doesn't mean in my terminology "literally achieving the worst thing", it's rather more like "optimizing" in that it's about progress in a direction. A longer way of saying it is "applying pessimization pressure".)
I realize applause lights around here are "capabilities bad" and "science good" and "conceptual progress good", but if you are urging more careful ethical and strategic reflection on the part of others, then perhaps consider examining this more carefully.
You're touching on an important point here, and upon reflection I think it's a deep gap that I'll need to bridge somehow (also tagging @Wei Dai). I do have some intuitions about why making conceptual progress/doing agent foundations does in fact differentially advance alignment over capabilities much more than all the other stuff. But the rea...
Ah, that clarifies things. (To be clear I'm pro science & conceptual progress, am not anti capabilities, and am skeptical of the alignment / capabilities distinction as presently theorized. I agree regarding philosophical progress and scientific understanding changing concepts.)
I'll leave aside "did alignment leaders act net negatively with respect to their original goals?" and focus on what the harms were and how they happened. As you say, Sam Altman and other AI leaders did a bunch of AI capabilities advancement, and in some ways benefitted from the alignment community, with respect to recruiting, funding, and narrative.
EA is mostly distinctive here in being a group of people who significantly disagreed with MIRI while still doing a bunch of interpretive labor and also having power, wealth, and ambition. (I blame utilitarianism for FTX but not for OpenAI.)
Basic factors that may have contributed:
a) Disagreement about crucial considerations
b) Interest in being personally important / powerful / rich
c) Inevitability intuitions as you discuss. (I'm undecided about how correct / useful such intuitions were. While there is a prima facie case against "do bad because someone would oth...
am not anti capabilities
I fed an archive of Jessica's LW writings to Gemini and asked it to explain this (I was also wondering about her recent tweets stating a similar position):
Her relative lack of fear regarding AI capabilities seems to stem from a deep appreciation for intelligence, complexity, and truth-seeking, which she weighs heavily against the desire to simply preserve the current human condition. She notes that she inclines toward "axiological cosmism" (the idea that there are higher forms of value that humans might not currently understand, but which superintelligences would likely pursue, as she explores in Why I am not a Theist).
Here is a more nuanced breakdown of the intuitions and arguments she uses to question the "capabilities are bad" consensus:
1. She is open to the idea that higher intelligence might yield better values
Rather than assuming an AI will inevitably pursue a "dumb" or arbitrary goal (like paperclips), she entertains the idea that values and intelligence are entangled. In her post "Why I am not a Theist," she suggests that if she had a vastly bigger brain, she would likely have new, better, and more well-informed intuitions about what is valuable. Fr
Rather, it's that the explicit goal of the alignment community was to differentially advance alignment over capabilities, but instead they ended up advancing capabilities much more effectively than anybody else, while not advancing alignment much. This is a failure to follow the stated goal of colossal proportions, we in fact optimized the opposite of the goal. Why did this happen
My baseline hypothesis is: "because it was much easier to advance capabilities than alignment (especially when capabilities were weak)". And also: "Drawing more people's attention to the importance of AGI will inevitably cause some of them to race towards it". (Either because they weren't convinced of the safety part, or because they think that 'advance capabilities to win the race and get more influence later' is a good strategy for mitigating the safety risks. I think in practice the AI company founders are selected to be a mix of those two.)
This hypothesis seems really important for me, because if it's true, it's not clear that mistakes were made. Because it's not clear that there were alternative feasible routes which would achieved a significantly better ratio of alignment:capabilities progress. (The ...
i'm sympathetic to the critique of most modern ai safety work being terrible, but i don't understand why you think the agent foundations stuff (Lobian cooperation, Definability of Truth) is so good. it seems likely that intelligence is complicated enough that you can't have good physics-like theories, and so all of the theories have to have really fucked up spherical cow assumptions. even studying human intelligence more closely feels more obviously relevant.
it seems likely that intelligence is complicated enough that you can't have good physics-like theories
Two responses. Firstly: before Newton, would you have been saying that physics is complicated enough that you can't have theories as simple and powerful as F=ma? Before Darwin, would you have been saying that life is complicated enough that you can't have theories as simple and powerful as evolution? Before Turing, would you have been saying that mathematical cognition is complication enough that you can't have theories as simple and powerful as Turing machines? Before Pascal and Fermat and von Neumann, would you have been saying that rationality is complicated enough that you can't have theories as simple and powerful as expected utility maximization? Before Adam Smith... (etc).
In other words: the greatest scientific insights come from theories that are so elegant that it's unimaginable to almost anyone that the world could be compressed to that extent. And so it's structurally hard to distinguish "this seems messy because I don't understand it" from "this is messy". But the latter seems to have been a terrible bet to make in many fields.
Not all fields, though—some remain uncompre...
We haven't seriously tried to have physics-like theories of intelligence! Only a few people (various academics, MIRI) have done something close to trying. Compare to how much effort has been put into optimizing deep learning.
Maybe physics was also confusing and weird until we understood it.
Counterpoint: Aristotelian physics was mostly right.
Aristotle’s physics is the correct approximation of Newtonian physics in a particular domain, which happens to be the domain where we, humanity, conduct our business. This domain is formed by objects in a spherically symmetric gravitational field (that of the Earth) immersed in a fluid (air or water) and the main celestial bodies visible from Earth.
For a student who has learned physics in a modern school it may sound strange to start physics by studying objects in a fluid. But for somebody who hasn’t it may sound strange not to: everything around us is immersed in a fluid. Aristotle’s physics is a highly nontrivial correct description of these phenomena, without mistakes, and consistent with Newtonian physics, in the same manner in which Newtonian physics is consistent with Einstein physics in its domain of validity.
Insta-upvote. Keep writing :-)
Want to push back a bit on the agent foundations part, as someone who got into it very early and came up with a bunch of stuff (e.g. the Lobian cooperation paper cites me for the main result). I don't think AF has much connection to the alignment of AIs that are being developed now. I think AF is "only" an extremely fun field of math/philosophy. Whether it deserves money/prestige/etc is a question for someone else. I just love doing it, and have a bit of allergy to overselling.
Had a nice conversation with Richard a few days ago which I will now try to summarize parts of, for those interested, including especially my future self who might want to pick up the thread later:
Whereas if you have some deontological constraints to obey, you can try to rationalize why some loophole doesn't REALLY count as violating the constraint, but maybe it's generally harder to do this?
You can rationalize by finding loopholes, but you can also rationalize by choosing different deontological constraints or virtues, or different interpretation of them. This seems like a somewhat serious problem? I think it's easy to find a group of people who will all say that it's really important to be high-integrity, but then who will strongly disagree about what being high-integrity means and maybe think that the other people in the group are acting in a low-integrity way. There's questions about honesty binds you to something like never lying, or to something more like always giving people an accurate impression of all important facts, or something else. People disagree a lot about whether there's deontological reasons to not advance AI capabilities on the current margin. Or whether there's deontological reasons to not eat meat.
It's reasonably easy to pick a convenient position here, because the whole debate about how to pick virtues or deontological principles is pretty hard to gro...
good points. But a counterpoint: If a community forms around a certain set of virtues and rules at time T, then that set becomes sticky and harder to change and therefore harder to twist into letting you do the power-seeking or cowardly or status-seeking thing later. For people who don't have any community or ethical identity, yes, they can choose from the menu the one that most helps them get status power etc. But then once they've chosen it, there are switching costs. By contrast if you are a consequentialist, you can simply convince yourself that actually you need to JOIN the AGI company, or that actually you need to ACCELERATE chip production, or whatever. The evidence about what'll have good vs. bad consequences is constantly changing, so if you change your mind for rationalizing reasons it can be disguised (to yourself and others) as a change of mind based on good evidence.
These and other mistakes are reflective of deeper irrationalities. One crucial pattern is what I call “jumping down the slippery slope”: viewing an outcome as so inevitable that it doesn’t matter much if you contribute to it, in a way which leads you to become a significant force pushing the world further and faster towards the "inevitable" outcome.
I agree that this pattern is real and that EAs and AI safety people such as myself have engaged in it too much historically. But how far do you think we should go in avoiding this failure mode? Should we refuse to use ChatGPT? Should we refuse to talk publicly about superintelligence, and instead dismiss it all as hype, in the hope of bursting the AI bubble? (There are probably many thousands of people currently taking this strategy so it's not just a hypothetical!) My current answer to both questions is no and I expect you'll agree, so there's a question of where to draw the line. I'm interested in ideas for principled and/or historically validated answers.
My knee-jerk reaction here is "have you tried asking a five year old?". Spelling it out in more words: for most of these questions there is an obvious Right Thing To Do. Say true things, spread true important things, don't lie or strategically hide your views, don't accelerate capabilities, don't found capabilities companies, don't bullshit. Just normal five-year-old level ethics.
And look, I am explicitly on the record saying Value Judgements Are Usually Bullshit and Human Values ≠ Goodness and just generally opposing mainstream morality. But even I can look at the track record of EA and AI Safety and say "look guys, you have not been outperforming the Obvious Right Thing To Do with respect to AI, you should stop with the gigabrain takes and just do the Obvious Right Things".
The two paragraphs after the one you quoted are intended to give some guidance on this. Some elaboration on key points, moving from most to least actionable (many of which I expect you're aware of, but this is just what I have off the top of my head):
Thanks, great post, and I'm looking forward to reading the rest of the sequence.
But I think many of the readers that you're looking to convince will be coming in thinking that it is good for the world that Anthropic exists. Yes, capabilities were accelerated. But some high-quality alignment work has been done, at least more than would have been done by the counterfactual frontier AI company by the time ai was this capable. And now there's a frontier AI company that is deeply concerned about AI takeover risk and will likely support efforts to pause.
So I would be interested to hear more about why you think things are worse than they would be in a counterfactual where AI safety people did not push forward AI capabilities and join frontier AI companies.
And yeah, my impression is that not much has come out of agent foundations, which feeds into the counterfactual
There's a loop that consequentialists can get into that's something like:
The problem here is that you're not cashing out "the world being improved" in anything foundational, but rather in large part in the fact that you believe in yourself. When I think about Anthropic considering how they should orient to a pause, for example, I imagine them thinking "but actually we're so altruistic that it's bad for the world if we lose ground, so we shouldn't pay significant costs in supporting a pause". And so the impact gets pushed off for another cycle...
What is foundational enough that we can trust it to ground claims about who should have a lot of power? For instance, if Anthropic had made some scientific breakthroughs regarding alignment, then I'd feel significantly better about their existence. This is why I talk so much about science in this post, because I think that if you use a scientific lens then there's very little "high-quality alignment ...
If you looked at non-consequentialist ideologies that was contemporaneous with Bentham, so many of them advocated for policy positions that we now consider evil.
Nor are non-ideological people immune. Plenty of "ordinary men" do evil things when people around them do.
Thomas Carlyle coined the term "dismal science" to argue in favor of slavery, against people like Mill and other economics-minded utilitarians.
“AI alignment” has been shaped by OpenPhil’s allocation of money, as well as the intellectual influence of a cluster of people associated with them
Yes. The distortionary effects on AI alignment's development, and the conflicts of interest that thereby arise, from having one major funder are very strong. I suspect this will remain true in the next philanthropy windfall (if it happens).
FWIW I think it's less about having one funder, and more about having one funder with such a strong agenda that's in so much conflict with the agenda of the founders of the field.
I'm not sure I buy the deep implicit argument here that fundamental scientific understanding has generally had a robustly positive effect on humanity. Or like I buy the case for scientific understanding construed very broadly but I worry it's generalizing too widely. For any given (major!) scientific advance maybe I'm still only at 55-45 or at best 60-40 that it's good to have been developed at that time (or earlier) as opposed to later, whereas maybe you're at more like 90-10 or higher? Like I don't see why it largely supersedes tinkering/engineering at the margins.
Concretely maybe I'd say that atomic physics, say, or the mathematical physical theory of diffraction, was clearly not robustly positive for either humanity or the purported goals of the specific people who developed and funded the relevant fundamental breakthroughs, whereas the tinkering by Alexander Fleming and successors on pennicilin, the gradual invention of paper in China, the printing press in Europe, experiments in variolation, and the experiments with crop yields by Borlaug plausibly had robustly positive effects (at least at timescales similar to the timescales that fundamental research gets judged by[1]).
o
Thanks for your recapitulation. I'm not very familiar with the early history of the field and found it quite informative.
But making AGI “open” was close enough to the opposite of what early rationalists wanted that it was clear that something had gone badly wrong.
I've just looked a bit through court documents of the Musk v. Altman lawsuit and found some notable exchanges but will limit it to two examples:
1) Demis Hassabis confronted Elon Musk about their open source vision for AGI, using a reference to the same article of Scott Alexander that you linked in your retrospective. Ilia Sutskever already commented it already back then in 2016 with "As we get closer to building AI, it will make sense to start being less open."
2) Brockman's journal contains this section from 2017:
...- gdb: if you and sam get more involved, will be unstoppable.
- ilya: new structure will help a lot too.
- elon: alright sounds good. game is afoot. gonna be battle.
- gdb: pretty soon will remember the days when enemy was just DM.
- elon: hah. don't want to pave road to hell with good intentions.
- ilya: don't create the AGI before making it act in our best interests. should be fundamental tenant. especially once we h
Strong upvote, as I consider any effort towards agent foundations heroic prima facie. However, two objections:
(i) Science generalizes further than engineering, thus any scientific insight is more capabilities-counterfactual.
For example, Legg & Hutter heavily draw upon the (albeit controversial and relatively primitive) science of g in their works on understanding intelligence (cf. Universal Intelligence, Chapter 2.1). This work arguably made it possible to even understand what AGI is in the first place (note: It should be said that Goertzel doesn't seem to believe in g).
(ii) Science is easier to drown out in noise with funding and prestige.
For example, the foundational works of Turing (1936) and Post (1936) received almost no attention upon publication. Sudan (1927) showed the Ackermann function is computable but not primitive-recursive, not recognized until 1979.
While the principle of "two formalizations agree (e.g. Church-Turing thesis)" makes some automation possible, obviously, if multiple formalizations agree, we still have the legibility problem.
That is, two of the key hard problems in alignment (capabilities work by accident, lack of legibility mechanisms) strike even ...
Some other reasons why a more engineering mindset was adopted in AI alignment as opposed to an idealized insight focused path that is more positive than the reasons you brought up is:
Science is partly a coupled process between engineering and theory, it is also hard to do causal inference from what counts as theory and what counts as practice.
Would you say that the discovery of the higgs boson was something that happened as a consequence through theoretical or practical physics?
What about deception, inner misalignment, general interpretability methods, if you trace their intelluectual lineage where do they come from? They're not fully agent foundations but if you compare if they're more ML based or coming from the larger space of theoretical alignment research I would attribute more causal influence to theoretical AI Safety. This is what agent foundations was up to like 3 years ago!
Or what do we mean by agent foundations here? What would you draw the boundaries around? Is it maybe better to use the word Theoretical AI Safety research?
Under slower takeoffs, engineering mindset is fine, because you have more hopes on fixing the problem iteratively, and this is due to the fact that you can assume that AI capabilities are more bounded than thought
Given specific assumptions about scientific progress where we can iteratively improve it and it is clear how we would ...
I would have much less beef with the engineering mindset if it hadn't accelerated capabilities so much, as I'll explain in the next post (and also if it weren't so related to conceptually confused ML research, as I'll explain in the post after that).
My main issue is The Counterfactual Quiet AGI Timeline. Scaling laws of neural nets had capabilities become more like treasures waiting for the right amount of compute. Once anyone invested the compute, the treasure would be his and everyone would rush for similar treasures, until one of them summons demons...
What would be a possible agent-foundations insight useful for alignment but not capabilities?
Or would the strategy be "hope no one runs away with this to accelerate capabilities"?
Another example: Von-Neumann-Morgenstern/Behaviourism arguably gives you a "reward" as a meaningful concept; so plausibly accelerated RL
The phrase alignment community confuses me. How do different people here understand the meaning of the phrase? What would happen if we played Taboo with it? What would we say instead?
I have questions about what we mean by the alignment community:
My gut feel is the phrase alignment community probably glosses over too much. Using a few more words is worth it, I think.
For example, it seems tempting to ascribe agency and thus culpability to the community. But does saying "the community" help? We want to learn and course-correct; to do ...
I'm glad somebody is laying out this story from an "original MIRI-style AI safety" perspective.
I do think that mechanistic interpretability and technical AI safety has gotten surprisingly good in the past 2-3 years, and that the Amodei/Christiano worldview, while probably still subtly flawed, has been resoundingly vindicated by events, better than anybody else's predictions back then.
I also think that we all underrated the value of building LLM-ish things in the first place. It would have been safer to not go in this direction at all, but would it have been better? This is a unique (literally) opportunity for the sorts of people who are interested in AI, namely us, to affect the world substantially, in the directions we choose. It is also very interesting, and economically productive. Are you sure that's not worth it?
From my next post, that I'm currently finishing up:
The prosaic AGI intuition has been vindicated since then: we’re now much closer to building AGI, and we haven’t learned any fundamentally new things about intelligence in the process. But the reason I only called it a partial paradigm shift is that “prosaic AGI” was both a prediction and a research strategy, and the two are less closely-related than they seem.
I am confused about how to relate to the rest of your comment. E.g. saying "technical AI safety has gotten surprisingly good" seems pretty bizarre in the wake of all the autonomous hacking incidents—even lab folks aren't saying that any more.
Your last paragraph gives me such a visceral flinch reaction that it's hard for me to engage directly with it. I could say more about the flinch reaction if you want (though it'd involve me trying to form hypotheses about your internal state in a way that's rude by conventional norms).
oh, please do talk about the visceral flinch, feel free to form hypotheses about me.
I'm still learning, but some things that seem "real" to me in technical AI safety are like, computational mechanics, singular learning theory, the beginnings of some multi-agent game theory, some "what is an agent" theory with causal graphs, etc. it isn't at a stage to be implemented into AIs in practice yet.
I think the autonomous hacking incidents absolutely indicate a problem with the current state of AIs and probably also the organizational culture of (parts of) the labs. I'm not claiming "the AIs are already safe." And I don't believe in alignment by default. I just think there are research threads besides agent foundations that are fruitful.
some things that seem "real" to me in technical AI safety are like, computational mechanics, singular learning theory, the beginnings of some multi-agent game theory, some "what is an agent" theory with causal graphs, etc.
I like these lines of research too, but note that none of these are the kind of work that's been promoted under the banner of "prosaic alignment", or that requires any hands-on engagement with LLMs. And at least the latter two qualify as agent foundations in my book. So insofar as these are the best examples of technical AI safety progress I'd account that as weighing clearly against the prosaic alignment worldview, moderately for the "mainstream learning theory" worldview, and moderately for the agent foundations worldview (especially given how much more attention, talent and funding the former has received within the field of alignment than the latter two).
Re your last paragraph: it feels kinda like the thing that journalists do when they want to frame a narrative without actually pinning down concrete opinions or cruxes. I'm obviously not "sure" about such high-level questions, but the main substance of your paragraph seems to be "this direction is less safe bu...
One other thing I want to make more explicit, re: "But "we" are the people who are trying to make AI more safe, so this seems like a failure."
"We" are also a lot of other things. We are, more or less, literate intellectuals, in a society that is rapidly becoming hostile to such people and the views they hold. "We" appreciate science and technology. "We" are interested in minds and understanding how minds work through computation, because we like having minds and understand that computation was designed in the first place to model thought. That's why AI is a convergently appealing idea in the first place, and why the original AI safety people all started out as AI enthusiasts.
LLMs strike me as a literally unique opportunity for "mind-enthusiasts" to influence the world with our values, in the 2020s. (I get the sense that, despite our differences, you and I have many values in common, in the usual directions that old-school LessWrong and its antecedent mailing lists differ from the rest of the world).
That's a main crux for me, with regard to the question "would it have been better if DeepMind and OpenAI and Anthropic hadn't developed & scaled LLMs in the first place?" I'm comin...
A lot of my 2010s writing was literally me parrotting other people. It's fair to hold me "responsible" for it, because I did write it, but I don't think it was authentically mine, and I don't agree with it now. Right now, I may sound dumber or less brave, but my words are more aligned with my actual own thought processes and actions.
In person, tbh, I expressed values I've had all along. I'm a perfectly normal central-tendency libertarian who hasn't changed a political opinion since 2010. I also have some grounded life experience in what healthy relationships & mental attitudes are, and what sorts of framings are indicative of an unwholesome/counterproductive attitude, & i might have pushed back on you there in a way I wouldn't have when I was younger & more timid.
Of course if it would literally kill everyone I will (with some reluctance) admit we should give up wealth and influence and cool intellectual fun and short term public benefit. I have profound uncertainty about how we should be thinking about "will it literally kill everyone" and I am hedging my career bets to only do things that seem good in a variety of scenarios.
Perhaps it is a mistake to speak up at all ...
I also think that we all underrated the value of building LLM-ish things in the first place. It would have been safer to not go in this direction at all, but would it have been better?
My answer is "yes, it'd have been better", because this sort of thing would have happened sooner or later. Because it's happening sooner, we have less "serial research time" for deep thinking.
All the interesting economic productive stuff still happens eventually. The part where "now it is concrete, and easier to think about, for people who benefit a lot from seeing concrete things in front of them" would have happened eventually. (Both for researcher-types and politician types). It could have happened in a world where we'd made more substantial strides on agentfoundationsy stuf.
I'm still learning, but some things that seem "real" to me in technical AI safety are like, computational mechanics, singular learning theory, the beginnings of some multi-agent game theory, some "what is an agent" theory with causal graphs, etc. it isn't at a stage to be implemented into AIs in practice yet.
I think there's a "." missing somewhere here? (This looks like a list of things, the first few of which are supposed to b...
Thanks!
- First, naive extrapolation of trends points towards RSI-capable AGI within the next decade, with pretty decent odds on "within a few years" if not stopped by something. (This is me mostly deferring to the AI Futures team, as well as staring that the METR graph and subjective experience at seeing AI improve over time)
AFAICT, the graphs show AIs getting better at the kinds of things they're good at, while not showing the ways in which AIs continue to be stalled and in which they've shown little signs of improving for the whole time we've seen them. (Analogy: if you see graphs where solar panel efficiency keeps improving, that lets you extrapolate that they're going to be increasingly efficient at producing energy, but it doesn't tell you much about their potential for anything else.)
As Rob points out about the METR graph in the article that you linked in the next paragraph:
...Firstly, this isn’t all tasks. It’s fairly cleanly specified software engineering, cyber, and machine learning tasks. And that’s probably the single thing AI is best at in the entire world. The reason it’s so good is that coding is a domain with extremely good feedback density — that is to say, lots of feedb
You think “EAs should’ve been more cautious about waking up various external groups like the public, academia, industry, governments” and “EAs should’ve been less power-seeking because other people will notice those power-seeking patterns and there will be a backlash”. Maybe there’s a tension there.
It’s useful to contrast prestige-orientation with Eliezer’s alternative recruitment strategy—writing Harry Potter and the Methods of Rationality—which was closer to a prestige-minimizing move, yet which was much more successful in recruiting people who could think clearly about alignment.
This is interesting. I wonder if more like that could be done and if it would be better.
As late as 2024, Yann LeCun was still declaring that LLMs “can not solve problems they haven’t been trained on”.
His views on the topic are a rock with 'LLMs cannot learn/scale/generalize' written on it. For example, from a few months ago:
...LeCun defines intelligence as “the ability to accomplish new tasks you’ve never been exposed to and solve new problems without any prior training” — a skill he says LLMs do not possess. He argued that because of this, LLMs will not be able to reach human intelligence, saying that “human-level AI will require real world dat
Excited for this!
...This kind of “engineering” mentality contrasts sharply with Eliezer’s original vision of alignment as the development of a powerful new scientific paradigm—e.g. see this post comparing agent foundations to Newtonian mechanics. It’s easy to make arguments on a case-by-case basis for why engineering work might be good for the world. However, the field of alignment is explicitly trying to do work that has predictably beneficial effects on an unprecedentedly large, world-historic transition. If we didn’t have such clear examples of scientific
and first (to my knowledge) made the link to a limiting infinite-width Gaussian process (which later evolved into Neural Tangent Kernel work.)
This is true. Not only that, but the use of Gaussian processes in machine learning comes from Radford Neal's thesis and that they're the limiting behavior of wide NNs in the first place.
Many fields of science were infeasible to develop until the corresponding engineering work had already been performed. The canonical example is the steam engine, which predates thermodynamics by over a century. Those early, commercially implemented steam engines allowed the scientists of the day to benchmark and measure the properties of actual examples of their field of study. There would be no sense in castigating the scientists of the day for not having worked enough on theory prior to the engineering work. The downside of this reality, the inefficiency...
The canonical example is the steam engine, which predates thermodynamics by over a century.
That's not true. The Watt steam engine was made possible by the thermodynamic research of Joseph Black, Watt's friend and financial patron, who discovered heat capacity and latent heat, thereby making Watt's separate condenser possible. Savery's earlier engine was based on Papin's research. etc.
It's true that thermodynamic science depended substantially on engineering progress as well, and worth noting that Black and Papin were more like experimentalists than like today's idea of a theorist who works with symbols and abstractions, so your point is correct even if your history is wrong.
To be clear, I’m not taking a strong stance in this sequence on whether AI will go well or badly—that seems up for grabs. My concern is that the alignment community had a plan to make good outcomes more likely (differentially advancing alignment over capabilities) but has mostly pushed the world in the opposite direction, while some parts of it gained a lot of power by doing so.
I should note here that the original goal mentioned in the paper here was to create AIs with desirable goals, and the plan you have outlined is an sub-goal to make the original succ...
(I don’t know how intentionally MIRI facilitated these kinds of interactions; I also don’t know how Elon first got involved.)
In 2010 when Thiel met Hassabis (probably at the Singularity Summit?), MIRI was still the Singularity Institute for Artificial Intelligence; Eliezer had pivoted away from transhumanist accelerationism towards FAI theory, but the institution as a whole had not yet done so.
...Even when the research wasn’t that fundamental, it was able to grapple with ideas that other intellectual communities simply weren’t able to collectively think about. Consider Omohundro’s paper on convergent instrumental goals, or Eliezer’s paper on intelligence explosion microeconomics. Neither of these contain powerful or surprising results—they’re just fairly straightforward analyses of concepts that can be explained in a single sentence. But no other community was able to reliably produce or build on such analyses. This effect is even starker when thin
I wish that the sequence was there to read. However, even this would benefit from doublechecking for mistakes like mocking people for trying to ban open-source models. The case against open-sourced models is that they lack safeguards unless they are as capable of goal-guarding as Agent-4, thus allowing terrorists to use open-source models to hack into important systems or to create bioweapons.
The framework also has a more severe issue. Your most recent quick take mentioned that "AI governance would have been less likely to throw in with the Democrats in a ...
You allude to self-deceptive reasoning in the introduction; What's your perspective on the extent to which prominent figures in AI safety (both now and historically) who have acted in ways you describe as power-seeking were doing so consciously vs subconsciously? This feels like a pretty load-bearing distinction, especially when it comes to thinking about how to intervene to make things better (conscious => get better at trusting the right people and put prominent figures under more scrutiny, unconscious => that + much more self-reflection amongst other things).
A central example is Sam Altman’s original email to Elon about founding OpenAI: “Been thinking a lot about whether it's possible to stop humanity from developing AI. I think the answer is almost definitely not. If it's going to happen anyway, it seems like it would be good for someone other than Google to do it first.”
I disagree that this is an example of a mistake, by Sam Altman seeing a bad outcome and doing the wrong thing to mitigate it - instead, it seems to be more accurate to model it as seeing an opportunity for a chance to push some buttons to climb a power ladder and doing so.
Great post, thank you very much for writing this.
I'd be interested in your thoughts on DCI, particularly any criticism or concerns you might have about it. I'm intending on drafting a longer post on its philosophy and theory of change, but the rough summary is that I see it as a potential way of paradigmatizing the science of generalisaton.
I'd also be interested in your thoughts on the work being done by Resolution, and their hopes around understanding LLMs.
Your distinction between scientific progress and engineering, made me think that there may be an important feedback loop missing from this picture.
One limitation of purely theoretical prediction is that we tend to consider only these scenarios, which we are able to imagine. Meanwhile, systems operating in real-world generate surprising behaviors. In recent months, we have seen, among other things, spontaneous workarounds of the ethical constrains, the emergence of unexpected forms of coordination, and cases in which pursuing a goal took priority over ethic...
This sequence is about the last decade in AI alignment. Over five posts, it recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progress in favor of iteratively improving existing systems and attempting to gain technological and political power. I also describe (in subsequent posts, which I'll upload over the next few weeks) how fear and (self-)deceptive reasoning made the field one of the biggest forces pushing AI capabilities forward over the last decade, especially via significant contributions to the scaling of LLMs and the development of ChatGPT.
Zooming out further: the two leading AGI companies, which are locked in an intense rivalry, were both explicitly founded under the banner of AI alignment, and got off the ground in significant part due to alignment-oriented ideas, talent and resources. People in the field often sense that something must have gone wrong to get here, but don’t know how to allocate responsibility (aside from blaming Sam Altman and sometimes Elon), and fall back on assuming that “the ship has already sailed”. But in this sequence I characterize our current situation as resulting from a pattern of mistakes which is both recognizable in the past and actively ongoing, and which if continued will cause similar kinds of dysfunction over the next decade.
To be clear, I’m not taking a strong stance in this sequence on whether AI will go well or badly—that seems up for grabs. My concern is that the alignment community had a plan to make good outcomes more likely (differentially advancing alignment over capabilities) but has mostly pushed the world in the opposite direction, while some parts of it gained a lot of power by doing so. This is not trustworthy behavior, and should be a big update about how well the community will use its power going forward. In particular, it’s very bad for the world that the community doing the most to steer the future of AI isn’t really trying to distinguish the extent to which its leaders are sincere vs sycophantic vs power-seeking.
I care about this significantly more than I care about the object-level effects of accelerating capabilities, because the integrity and rationality of a few key decision-makers will likely shape the coming decades. And yes, there are others with power over AI (like Sam and Elon) who have less integrity in most ways than alignment leaders. However, in my mind the level of adversarial dynamics within the field makes transparency, integrity and accountability more important rather than less: lacking integrity makes you much easier to manipulate, as I recount in the post on Fear and Anticipatory Obedience.
Unfortunately, the alignment community is doing very badly at learning from the past decade, or holding anyone accountable. Indeed, it’s pursuing many strategies which seem likely to recapitulate previous mistakes. Four of the most prominent, which I’ll discuss in the final post, are:
These and other mistakes are reflective of deeper irrationalities. One crucial pattern is what I call “jumping down the slippery slope”: viewing an outcome as so inevitable that it doesn’t matter much if you contribute to it, in a way which leads you to become a significant force pushing the world further and faster towards the "inevitable" outcome. A central example is Sam Altman’s original email to Elon about founding OpenAI: “Been thinking a lot about whether it's possible to stop humanity from developing AI. I think the answer is almost definitely not. If it's going to happen anyway, it seems like it would be good for someone other than Google to do it first.”
Many strategies pursued by the alignment community (e.g. the four listed above) showcase this same pattern. You can view it as a result of individuals inappropriately reasoning about their marginal impact despite their actions often having extremely non-marginal effects (in part because so few others were, and are, taking superintelligence seriously). More fundamentally, you can view it as a result of individuals inappropriately reasoning about individual impact rather than thinking about the policies that they’d recommend for the field as a whole (which more sociology-style reasoning about norms and preference cascades, or FDT-style reasoning about entangled decision-making, would have prevented).
However, even given these conceptual errors, people wouldn’t jump down nearly as many slippery slopes if they weren’t driven by strong emotional instincts. For Sam, many of those instincts seem to be about accumulating power, from a perspective where nobody else can be deeply trusted. For the alignment community, people often hold narratives like “I need to save the world” or “I need to have impact soon”, and are scared enough of failing that they counterproductively narrow their vision (e.g. fear about “short timelines” gets in the way of pinning down what they’re even timelines to). Underneath that, though, there’s a similar (albeit weaker) kind of distrust in one’s relationship to the rest of the world, as I’ll detail in the final post.
Even if you find my explanations uncompelling, I hope that the abundance of detail I’ve included in the sequence helps you formulate alternative hypotheses about what’s going on. I’ve tried to be extremely transparent about what I’ve observed throughout my career, telling as many anecdotes as I can that give color on the events of the last decade. This involves being franker (and using more names) than is normal—in part because I strongly believe that people who try to significantly change the world (whether motivated by altruism or otherwise) are implicitly opting in to a high level of scrutiny. I’ve also been upfront about the many ways that I’ve personally failed. (Due to the sheer number of people discussed in this sequence, I haven’t run it past most of them before posting, and am open to corrections and/or additional anecdotes.)
As a final preamble, I recognize that this sequence is negative about many things. But I continue to believe that the alignment community is capable of more clarity and sincerity than any other similarly-sized intellectual community existing today. And I’m also feeling better on a personal level than I ever have. It’s very refreshing to pin down specific mistakes that led to specific failures, rather than living in a miasma of confusion about why bad things keep happening despite our best efforts. I think that the future of AI alignment—and with it, the future of humanity—is very much up for grabs. There are pathways hazily visible to me that (some subset of) this community could plausibly take towards extremely good outcomes—despite its flaws, people in it are trying to reason clearly and formulate large-scale plans to an extent that is extremely rare. The main thing blocking us is our inability to learn from our past mistakes.
I’ll start by laying out the intellectual approach that allowed early rationalists to think clearly about AGI even in the era of very narrow AIs (Conceptual Clarity and Scientific Progress). The second half of this post (Orienting Towards Prestige) covers early engagement between rationalists, Silicon Valley, and effective altruists, and the ways that backfired. The next post (Pragmatism and Pessimization) details how a small group of researchers nominally pursuing “prosaic alignment” were responsible for a huge amount of AI capabilities progress, and dramatically amplified the race dynamics between AGI companies. The third post (Conforming to the ML Community) explores attempts to recruit mainstream ML researchers to do alignment research, and the significant costs of doing so. The fourth post (Fear and Anticipatory Obedience) explains the dynamics which prevented people from speaking out about the failures they were seeing, especially at OpenAI. In the final post (Deja Vu), I discuss ways we might recapitulate these mistakes, and how to avoid them.
Conceptual Clarity and Scientific Progress
The early rationalist community was a beacon of intellectual clarity. In this section, I’ll talk about what that originally looked like, and why it was missing from academic machine learning (and academia in general). In the rest of this post and the next, I’ll talk about how the field of AI alignment gradually traded that clarity away as it grew—first via prestige-oriented recruitment efforts, and later via developing concepts and frameworks which prioritized conformity to the norms of mainstream machine learning over insightfulness.
The rationalist community drew its early members primarily from the transhumanist community (c.f. the Extropian and SL4 mailing lists), and the econ blogging community (c.f. Marginal Revolution and Overcoming Bias). After Yudkowsky split off from blogging at Overcoming Bias, LessWrong became an online hub of people who were doing very deep thinking. Wei Dai and Hal Finney were two of the earliest cryptocurrency pioneers. Robin Hanson was inventing prediction markets (alongside many other important concepts). A logic professor I talked to recently expressed that Christiano et al.’s 2013 paper Definability of Truth in Probabilistic Logic was a groundbreaking result that should be in every logic textbook. Scott Alexander’s applications of game-theoretic concepts to politics (as well as his pushback on wokeness) made him one of the most influential political thinkers of the last decade, particularly shaping the worldview of the emerging power center of Silicon Valley.
I consider Bostrom’s work on anthropics, Eliezer and Wei’s work on decision theory, Leverage’s theory of psychology, and the Lobian cooperation result to also contain very deep insights—though they haven’t yet been built upon in ways which make that depth obvious. And, of course, people were developing a set of ideas about AGI which would prove to be far more predictively powerful than standard ML frameworks. Eliezer and Robin’s debates raised many considerations that are still shaping our thinking about AI almost two decades later. Shane Legg, who coined the term AGI (and cofounded DeepMind) was an early LessWrong commenter. The idea of learned policies having goals of their own (separate from their training objectives) was such an important insight that it has now become hard to appreciate how novel it was. Any way you slice it, this was an enormous concentration of intellectual progress.
Even when the research wasn’t that fundamental, it was able to grapple with ideas that other intellectual communities simply weren’t able to collectively think about. Consider Omohundro’s paper on convergent instrumental goals, or Eliezer’s paper on intelligence explosion microeconomics. Neither of these contain powerful or surprising results—they’re just fairly straightforward analyses of concepts that can be explained in a single sentence. But no other community was able to reliably produce or build on such analyses. This effect is even starker when thinking about less technical work—like Bostrom’s Fable of the Dragon-Tyrant, or Astronomical Waste, or Hanson’s thoughts on signalling (later elaborated upon in The Elephant in the Brain). When I talk about intellectual clarity, a lot of what I’m talking about is the ability to take ideas that are actually very simple, internalize them, and then use them as building blocks to construct the next generation of ideas.
More generally, one touchstone I’ll be referring to throughout this sequence is the idea that scientific progress proceeds by developing insightful new concepts, which link together to form a whole new ontology that replaces the previous ontology. Kuhn, Feyerabend, Koestler, Chang and various other philosophers of science have described a range of past breakthroughs which fit this pattern. Importantly, this view of science isn’t prescriptive about how to develop new concepts—it can be done via naturalist, experimental, mathematical, philosophical, or even mystical thinking. The quality of such work is often hard to evaluate at the time, but hindsight makes it easier to see who was aiming towards conceptual breakthroughs. And sometimes people are explicit about not doing so—e.g. one of the most senior alignment researchers at Anthropic recently told me that the best way for me to track if they were making progress on alignment was by using Claude and seeing how aligned it was.
This kind of “engineering” mentality contrasts sharply with Eliezer’s original vision of alignment as the development of a powerful new scientific paradigm—e.g. see this post comparing agent foundations to Newtonian mechanics. It’s easy to make arguments on a case-by-case basis for why engineering work might be good for the world. However, the field of alignment is explicitly trying to do work that has predictably beneficial effects on an unprecedentedly large, world-historic transition. If we didn’t have such clear examples of scientific theories generalizing extremely far, then this would be a very speculative strategy. So if you’re doing not-very-scientific alignment research with the aim of aligning superintelligence, you should expect most of your impact on the world to come from unpredictable higher-order effects of your actions (or predictable effects which you mentally blocked from consideration, as I describe in the post on Pragmatism and Pessimization). This problem is exacerbated if you backchain from alignment research going well to justify other kinds of work (like recruiting, communications, political manoeuvering, etc), since that introduces further complicated (and often adversarial) multi-agent dynamics—as I describe in the post on Fear and Anticipatory Obedience.
Unfortunately, most alignment research is no longer even aiming towards the kind of scientific progress I describe above. Agent foundations is the only subfield of alignment which consistently does so, and therefore the only one which I consider reliably good to generically promote.[1] Some parts of mechanistic interpretability are also building the kinds of understanding that could lead to a scientific revolution, but unfortunately they’re not very clearly-demarcated from the parts that might have large effects in other ways (like advancing capabilities), so overall I expect that field’s effect on the world to depend sensitively on the judgement and virtue of the individuals involved.
I’ll also briefly note that similar problems apply to most AI governance interventions, which are even more prone to backfiring (since modern politics is so adversarial). Even pausing AI progress, which could be extremely good, could easily be implemented in very bad ways—and almost nobody is thinking clearly about the differences between those. So the only outcome in the AI governance space that I consider reliable enough to backchain from is building (justified) trust between key actors—like different AGI companies, or the US and China. (Meanwhile cyberdefense and biodefense are robust in some ways—hence Vitalik’s advocacy for d/acc—but still require good judgement to do well. E.g. it’s easy for people in either field to reason their way into doing gain-of-function work, trying to ban open-source models, etc.[2])
One reason people are confused about AI alignment losing its ability to make scientific progress is that machine learning as a whole is also not a very scientific field by the standard I’m applying. Even most early AI researchers were more focused on building artificial intelligence than on understanding scientific principles of cognition. The rise of deep learning exacerbated this problem, as throwing more compute and engineering effort at an AI became arbitrarily scalable. In an important sense, the field of alignment is necessary because the field of ML didn’t prioritize gaining a deep understanding of the systems it was building. (Eliezer makes a similar point in this dialogue.)
A lot of the blame should fall on misguided narratives (common across academia) about what makes science work, which have been entrenched by the best-funded scientific institutions. A core scientific norm is that disputes should be resolved with reference to concrete empirical tests or rigorous proofs, judged by the scrutiny of one’s scientific peers. But that’s very different from the idea that ideas should be developed via paper-sized units of work which are each individually defended and justified. Historically speaking, the latter simply isn’t how the best science happened—Newton and Smith and Darwin developed their ideas via writing books and letters rather than peer-reviewed papers (and even Einstein didn’t encounter peer review until decades after his main breakthroughs). But the requirement to “publish or perish” is now so entrenched across academia that it produces strong streetlight effects.
Some concrete examples from ML: until recently, almost all RL theory focused on the unrealistically simple tabular setting, because it was easier to prove things about. I expect that there are important theoretical insights to be discovered about non-tabular RL, but progress towards them would require grappling with qualitative and fuzzy ideas for extended periods. Meanwhile, statistical learning theory spent decades focusing on the underparameterization regime, which doesn’t do much to explain generalization in neural networks (or biological brains). I don’t have enough context to give a confident explanation for the emphasis on underparameterization, but the ease of proving things about this regime seems like an important component.[3] In this 1995 commentary on NIPS (now NeurIPS), a statistician frustrated by how “everyone wants to be a theorist” writes that “mathematical theory is not critical to the development of machine learning. But scientific inquiry is.” He characterizes scientific inquiry as “sensible and intelligent efforts to understand what is going on”, and gives overparameterization of neural networks as a central example of a good target for scientific inquiry. Yet only after the rise of deep learning did phenomena like grokking, deep double descent, and memorization of random labels render this omission too blatant to ignore (though I’m uncertain about how much real progress subsequent theoretical work has made).
So the kind of conceptual thinking that I’m praising in the rationalist community is what I’d call the generative part of science, which elsewhere has been swamped by overly-zealous discriminative classification.[4] (See also Strevens’ insightful analogy of science as a coral reef.) Zealous evaluation also serves to entrench the power of existing academic hierarchies. As a case study, it's instructive to consider how the academic ML community oriented towards the concept of AGI overall. In some ways, AGI is a very simple concept: AIs that can generalize to a comparable extent as humans. Yet what we saw when the AI alignment community interacted with the academic ML community was something akin to an immune system response. Most scoffed—like Andrew Ng, who claimed that worrying about AGI risk was “like worrying about overpopulation on Mars”. Even when they did respond, it was with transparently bad arguments.[5] When I joined DeepMind in 2018, most of the researchers I met there still thought of the company’s own AGI-related mission statement as an eccentricity. As late as 2024, Yann LeCun was still declaring that LLMs “can not solve problems they haven’t been trained on”. For more on these dynamics, see Chapter 6 (“The Not-So-Great AI Debate”) of Stuart Russell’s Human Compatible.
Having said that, the field of ML was also reacting in part to the alignment community’s lack of appropriate discrimination. In particular, rationalists often treated informal, abstract arguments about AI risk as far more decisive than was warranted, in part due to an epistemology which claimed to supersede standard scientific epistemology (and in part due to a strong emotional orientation towards “saving the world”). Rather than focusing on further developing and clarifying its insights about AGI risk, though, the rationalist community spent significant effort winning its skeptics over, with largely regrettable effects. I think of this process in terms of three waves: Silicon Valley, the ML community, and the US government. I’ll discuss the first below, the second in a later post, and save the last for the final post in this sequence.
Orienting Towards Prestige
The rationalist community wasn’t disjoint from conventional prestige networks—for example, Hanson and Bostrom were professors.[6] Jaan Tallinn was around from pretty early on; so was Peter Thiel, who was introduced to Demis Hassabis and Shane Legg by Eliezer Yudkowsky, at an event cohosted by Thiel and MIRI.[7] Based on an interview with Thiel, Sebastian Mallaby claims in his book The Infinity Machine that "Eliezer Yudkowsky’s endorsement meant a lot. Thiel had known Yudkowsky for half a dozen years, and DeepMind was the first company that he had recommended." (Though note that Luke Nosek rather than Thiel ended up as the driving force behind Founders Fund's initial investment in DeepMind, according to Mallaby.)
However, attempts to recruit elites gradually became more publicly visible. Bostrom’s Superintelligence was a (NYT-bestselling) attempt to make AGI risk a prestigious concern, with an endorsement on the cover from Bill Gates. Various conferences organized by the Future of Life Institute collected growing numbers of notable figures (including Elon, though I don't know how he originally got involved). These efforts weren’t necessarily targeted specifically at Silicon Valley elites, but those were the main ones who took the ideas seriously enough to act on them. I’d count Dustin Moskovitz as another Silicon Valley elite who gradually became serious about AGI risk (with consequences that I’ll discuss at the end of this section).
I wasn’t present enough in the community at the time to have a sense of how explicitly people were reasoning about the value of outreach to prestigious elites. At the very least there was an implicit hypothesis that seemed straightforwardly plausible, which I’d gloss as “There are competent people out there in the world—look at the impressive companies they can build! We should recruit them as allies.”
But pretty quickly it became apparent that something was wrong with that hypothesis. For one thing, even very prestigious elites seemed less capable of sensibly discussing AGI than many anonymous commenters on LessWrong. A more dramatic datapoint came after Elon and Sam responded to concerns about AGI risk by launching OpenAI. I do think that there are some important and robust intuitions in favor of openness (e.g. hacker intuitions) which rationalists had been underrating. But making AGI “open” was close enough to the opposite of what early rationalists wanted that it was clear that something had gone badly wrong. In the past I’ve thought of the founding of OpenAI as an example of Silicon Valley’s extreme bias towards quickly taking action; now this seems absurdly charitable, and I think it’s better understood as a bias towards gaining power. To be clear, I hold Elon and Sam strongly morally culpable for this; I’m focusing on critiquing the alignment community instead because it seems more salvageable (though it’s also more morally culpable than Elon in e.g. its lack of political courage, as I discuss in the final post).
It’s useful to contrast prestige-orientation with Eliezer’s alternative recruitment strategy—writing Harry Potter and the Methods of Rationality—which was closer to a prestige-minimizing move, yet which was much more successful in recruiting people who could think clearly about alignment. To be clear, orienting towards prestige is not a bad thing in healthy social structures, where prestige correlates with competence, virtue and resources. However, being too focused on prestige makes you incapable of noticing when you’re deferring to unhealthy social structures. This kind of evidence takes time to accumulate, so I don’t blame MIRI much for reaching out to Silicon Valley elites early on; and even Superintelligence, insofar as it was a mistake, seems like a fairly understandable one. However, more EA-oriented people (especially those associated with OpenPhil) harmed the field significantly by conforming to existing power structures even when they should have known better, as I detail in the rest of this post.
The same year that OpenAI launched, Holden Karnofsky started to fund AI safety via Open Philanthropy. Holden had first heard MIRI’s arguments about AGI risk in 2007, but didn’t take them seriously due to MIRI’s lack of prestige. As he later recounted, he thought that “MIRI's lack of impressive endorsements from people with relevant-seeming expertise was the most important data point about it”; he was also influenced by “the general degree to which MIRI's views were seen as "wacky" and "silly" to a broad variety of people I spoke with”. On the object level, Holden also placed a lot of weight on the idea that AI would be a tool rather than an agent, and criticized MIRI for not taking that possibility seriously enough (you can read more of his engagement with MIRI ideas in this dialogue and this dialogue).
After the positive reception of Superintelligence by prestigious figures (including some ML researchers), Holden changed his mind, and gave OpenPhil’s first AI safety grant in 2015. However, despite spending 8 years being misled about AGI risk by over-indexing on prestige, he immediately directed the vast majority of his funding towards prestigious institutions rather than the rationalists who had laid out the case for AGI risk in the first place. While the importance of agent foundations research can be difficult to understand directly, MIRI’s prescience was clearly strong evidence that they had a deep understanding of the issue (plausibly too deep for Holden to appreciate), and any reasonable kind of hits-based giving would then have funded them to excess.
Instead, OpenPhil’s first grant to MIRI (in 2016) was only $500,000, and to a significant extent it was a “participation grant” to recompense MIRI for engaging with OpenPhil. By contrast, the previous year, Max Tegmark had received twice as much for the Future of Life Institute; and around the same time, Stuart Russell received 10x as much for CHAI. The following year, OpenPhil’s funding to MIRI was also less than 10% of their total “AI safety” funding—they gave $3.75 million to MIRI, almost $10 million to various university-affiliated groups, and $30 million to OpenAI (in exchange for a board seat for Holden). It seems reasonable to summarize this as Holden strongly calibrating donation size to conventional prestige.[8] More explicitly, Daniel Dewey’s main rationalization in 2017 for not giving MIRI much money was that agent foundations “has not gained much support among AI researchers”, and therefore wouldn’t be very useful for attracting new people to the field. In other words, he was making funding choices by deferring to the research taste of people who didn’t yet take AGI risk seriously.
Daniel ultimately concluded that “MIRI's current size seems to me to be approximately right”. Given how explosively the rest of the field was growing, this led MIRI to become a small, niche part of the field that it had founded.[9] I emphasize the relative sizes here because what we even consider to be “AI alignment” has been shaped by OpenPhil’s allocation of money, as well as the intellectual influence of a cluster of people associated with them. In particular, many of the mistakes outlined in my next two posts came from the version of alignment promoted by Holden, Dario and Paul. (This was both a professional and a personal clustering. OpenPhil’s writeup on its OpenAI grant ends with the following disclosure: “OpenAI researchers Dario Amodei and Paul Christiano are both technical advisors to Open Philanthropy and live in the same house as Holden. In addition, Holden is engaged to Dario’s sister Daniela.”)
Would the field have been redirected anyway by the sheer size and prestige of OpenAI? Perhaps, but OpenAI’s credibility as an authority on alignment depended in large part on the people who chose to associate with it. Without them, it would have been easier for the field to disown OpenAI’s approach to “safety”—though unfortunately even people unaffiliated with OpenAI were mostly too scared to actively oppose it, as I discuss at the end of the post on Fear and Anticipatory Obedience. It’s also important to note that, while $30 million is small compared to the billion dollars pledged to OpenAI when it launched, TechCrunch reports that only $133 million of that was actually donated. This would mean that OpenPhil’s $30 million was over 20% of the total charitable funding OpenAI ever received.
Trying to contribute a marginal 3% of OpenAI’s donations, and actually giving over 20%, is a great example of jumping down the slippery slope. But even aside from that, marginalist thinking about whether OpenAI would have redirected the field anyway is antithetical to upholding ethical standards. If two different groups are trying to do something bad, then the fact that it still would have happened if either had been removed doesn’t absolve each of responsibility—rather, it renders them both responsible for their participation in harmful group dynamics. In the next post, on Pragmatism and Pessimization, I’ll talk about other important ways that EA-style marginalist thinking boosted the development of AI capabilities.
Since I'm now an agent foundations researcher, you could view me as "talking my book" here. But it's worth noting that I was fairly dismissive of agent foundations during my first 5 years in AI safety—I instead worked as a research engineer, then as a PhD candidate in academic philosophy, then as a forecaster, then in AI governance. I picked up agent foundations only two years ago (despite lacking a strong mathematical background) due to the kinds of considerations articulated in this sequence.
I don't have a strong position that banning open-weight models is bad. However, while it mitigates risks from smaller actors (like terrorists), it seems likely to exacerbate risks from bigger actors (like companies or governments concentrating power via AI). I tend to be more concerned about the latter category—perhaps even in the context of biorisk. Government-sponsored gain-of-function research and military bioweapons programs seem pretty worrying, so banning open-weight models might disproportionately harm the development of defensive capabilities. Overall the main claim I'll defend is that banning open-weight models is not robustly good.
More generally, risks from omnicidal actors (like bioterrorism) seem sufficiently different to risks from power-seeking actors (like AI takeover) that I think we should mainly focus on the latter when doing strategic reasoning about the future of AI, and then primarily try to mitigate the former with domain-specific interventions (like biodefense).
Rif A. Saurous left a very helpful comment arguing against the claim I originally made in a draft (that underparemeterization was obviously not a good explanation for generalization in biological brains). Its details seems useful enough for the historical record that I reproduce it below in full:
I'll say some things I think I remember, partially jogged by Claude, but also admit this was close to 30 years ago, and frankly I'm still confused.
I was a graduate student in Poggio's lab at MIT from 1997-2002. We genuinely believed that Vapnik's learning theory was (in many ways) a "good explanation" for how and why biological brains worked. The Bayesian side in those days seemed similar --- for instance MacKay (who was very well-respected as a "real thinker") wrote about Bayesian Occam's razor, which is basically the same story.
I glibly phrased this as a story about "underparameterization", but it's not quite that. Our most powerful artifacts were SVMs with Gaussian kernels, which had a parameter per data point, and we knew those were our best performers. We also had theory results like "Boosting the margin" and "For valid generalization, the size of the weights is more important than the size of the network". Also, Breiman in 1995 wrote "Reflections after refereeing for NIPS", which I hadn't seen before but directly includes questions like "Why don't heavily parameterized neural networks overfit the data?" Also, Radford Neal wrote a (widely known) PhD thesis in 1996 on Bayesian NN's that argued against limiting network size, and first (to my knowledge) made the link to a limiting infinite-width Gaussian process (which later evolved into Neural Tangent Kernel work.)
We genuinely believed capacity control was key to generalization, but that's not quite the same as requiring underparameterization. And we did connect all this frequently to biology: Poggio's lab mixed learning theory folks (like me at the time) with computational neuroscience people, and we often cross-collaborated.
We certainly didn't have the modern insights from "Benign Overfitting in Linear Regression". Instead, we'd built (incorrect) insights from low-dimensional problems, where to fit a lot of data your functions have to oscillate wildly everywhere, whereas in high-dimensions you can hide the oscillations in dimensions where there's no data. We had early notions of intrinsic dimension and manifolds.
One thing I'll add that does look very bad in retrospect --- we more-or-less explicitly dismissed neural nets as "the way forward". We were all friendly with Yann LeCun, he'd come give talks, show us the cool results he was getting, but whenever we tried to replicate his work, we'd fail to train. An analogy is how biology often still doesn't replicate across labs; LeCun had "training taste" that we didn't know how to imitate. So we retreated to a joint package of "convex optimization is better because the theory is better and because it's easy to train", and we let those circularly reinforce each other, until accelerators came along ten years later and Hinton and Ilya and co. showed us how wrong we were.
I'd draw a similar link between analytic and continental philosophy. The former is extremely discriminative in the precision of the reasoning it accepts, without being able to generate creative new ideas—while the latter has the opposite problem.
Note that Chollet’s original title (still recorded in the URL) used "Impossibility" rather than "Implausibility".
An underappreciated fact is that Hanson was originally hired by Tyler Cowen, who thereby played a significant role in the formation of the rationalist community.
Per Mallaby, Legg recounts Yudkowsky saying "these are some of the smartest guys in the whole field of AI and they’re starting a really ambitious company." However, this seems to me like a (potentially quite lossy) summary rather than an attempt at an exact quote. Note that I didn't discuss this introduction in the original version of the post, because I hadn't yet read the relevant section of The Infinity Machine.
Eliezer writes about these dynamics (likely inspired by his interactions with OpenPhil) in this post.
Eliezer later wrote “I think it was a huge, huge mistake that more money was not spent on AGI alignment when it was small and weird and unproven. The resulting damage was not something that could be fixed by any or all of the money that became available later.” However, note that I’m not sure which period he was referring to, or who he thinks made that mistake.