Note: the modeling assumptions and conclusion are Thomas Kwa's opinion, and others at METR disagree. [1] Also, the math was checked by Claude but not a second human. Introduction Anthropic's RSI blog post reported that in Q2 2026, Anthropic contributors merged 8× as much code per day as in the...
In this post, I describe a simple model for forecasting when AI will automate AI development. It is based on the AI Futures model, but more understandable and robust, and has deliberately conservative assumptions. At current rates of compute growth and algorithmic progress, this model's median prediction is >99% automation...
This work was done while at METR. Introduction GDM recently released a paper (Emmons et al.) showing that, contrary to previous results, the chain-of-thought (CoT) of language models is more faithful when the model’s CoT is necessary for it to complete a task. They examine three settings where an actor...
Summary METR blogpost | X | Github In the paper Measuring AI Ability to Complete Long Software Tasks (Kwa & West et al. 2025), METR defined an AI model’s 50% time horizon as the length of tasks (measured by how long they take human professionals) that it can complete autonomously...
arXiv | project page | Authors: Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, Gao Huang This paper from Tsinghua find that RL on verifiable rewards (RLVR) just increases the frequency at which capabilities are sampled, rather than giving a base model new capabilities....
As Americans know, the electoral college gives disproportionate influence to swing states, which means a vote in the extremely blue state of California was basically wasted in the 2024 election, as are votes in extremely red states like Texas, Oklahoma, and Louisiana. State legislatures have the Constitutional power to assign...