I'm using Astra for side projects, and having to go back to Claude for work is annoying. I didn't realize just how often Claude is wrong about stuff and needs to be corrected, until Astra just wasn't wrong.
The personality is also great. Astra never tries to "push back" or have a personality. It just does the thing you asked. It still suffers from the same high level reasoning blindness as other models, perhaps a little moreso, so you have to be its strategizer/planner/manager. But it handles all the little tasks you want to give it, and makes any kind of computer project much easier. I find myself not needing to bother verifying its work. It's very good at verifying it's own work.
AGI is here.
I tend to call it “proto-AGI”.
It’s quite close. The details of its ARC-AGI-3 performance are very impressive (not just the score, but what it was doing to achieve that).
Also I had a not-very-well-known open math problem from the theory of complete lattices and Scott topology for the last 30 years (not too difficult I think, but I was not able to solve it despite many repeated attempts or to convince technically stronger people to invest enough effort). I started to give it to models since last Summer, and they gradually have gone from being quite useless and incompetent to being helpful and showing promising ways and lines of atrack and formulating useful correct lemmas. Finally, Astra (non-Pro) has solved most of it in 10 min of thinking from a one-shot simple prompt (it did have access to earlier conversations in my account, and there was an element of luck, as those conversation led the model to a very recent paper, not directly related, but containing some useful material; the solution was very elegant; still there is remaining work to fully verify and present well and so on; but for the purpose of model evaluation I am inclined to score this one as “done”).
Astra impresses on the Epoch Capabilities Index, scoring 169. This is directly on the OpenAI trend line here. Fable 5.1 was below trend and scored 163, same as Fable 5.
Now the two models score 167 and 164.
Astra is an excellent model. The jump from Sol to Astra is larger than the jump from Fable 5 to Fable 5.1. This is a big deal.
Astra is the best model for what one would broadly call ‘ambitious projects,’ and likely has the highest raw intelligence factor of any model. These are the largest jumps.
It is amazing at doing things in 3D, or anything involving games. Astra also excels at computer use, and at subagent coordination.
Many benchmarks show dramatic jumps from all previous models. Where Astra is good, it can be in a league of its own.
That does not mean Astra is in its own league across the board. Fable 5.1 is still a Claude. Astra is still a GPT. If you have a strong preference for one over the other, that still applies. For many purposes, especially involving back-and-forth discussions, Fable 5.1 is still my top choice. Fable remains my primary editor.
If you want the best answer to your questions, you should ask both models.
Regular coding is getting less of a focus. Astra is not a quantum leap there, but of course it is very good and makes progress over Sol.
This is the first time a debate over whether a model ‘was AGI’ felt non-silly. I do not think it is AGI, and I would warn against the dangers of using that label prematurely, but I would not laugh at you for disagreeing.
This is also a strange situation in that OpenAI has already soft announced that they have an internal model a level above Astra, as I will cover when I address Navier-Stokes.
My recommendation is that you use both Fable 5.1 and Astra on your most difficult questions, and experiment to see which things each one does best for you.
Table of Contents
Meanwhile
The backlog of things to discuss is not getting smaller, so I will take this moment to encourage everyone to read this excellent essay by Dario Amodei, We Must Pace the Frontier.
This is the core thing he calls for, including a unilateral commitment:
I will have full coverage of that next week. Along with that essay, the current queue includes at least that, Navier-Stokes and Astra-2, Anthropic’s Misalignment Report, Anthropic’s Countering Misuse, a thinkpiece on personal AI and the law, a post called Claude Talk, and Fable 5.1 (and Astra?) Model Welfare.
Okay, back to Astra.
The Official Pitch
The pitch is that this is AGI.
Jensen Huang also calls it AGI, but he calls everything AGI.
The standard release video (3 min) has a stunning 125 million views on Twitter.
Here is the pitch from Astra itself, according to Pangram:
As is often the case, scientific discovery was highlighted.
An early demo Noam cites was Astra proving there are infinitely many pairs of consecutive primes whose distance is at most 186, down from the previous bound of 212 (also from OpenAI), although there is some dispute over what exactly was or wasn’t proven here.
They show off a variety of cool demos and capabilities: Acing the Financial Modeling World Cup and navigating spreadsheets at 4x human speed, doing PCB in KiCad, creating a 3d model of a car transmission, filling out a 1040 and so on. They demo ordinary tasks. It’s all cool, but we lack comparison points.
The professional work pitch is that Astra handles complex tasks and adheres closely to templates and instructions, especially when creating presentations. It can translate images into identical-looking spreadsheets, yay.
They highlight Blender 3D models, which many others were also impressed by, see the section In 3D. Game creation is also confirmed as super impressive.
Our Price Cheap
Astra is a premium model.
The headline price is $10/$50 per million input and output tokens. Cache writes are $12.50 and cached input only $1.
If your prompt has more than 272k input tokens, prices on input double, and prices on output go up 50%.
Fast mode costs double.
Fable 5.1 has the same $10/$50 headline price, but its cache reads are only $0.25.
Practical costs come down to token efficiency. Artificial Analysis thinks Astra is only about 60% more expensive than Sol in practice, and that it is a lot cheaper than Fable 5.1. They used Fable 5.1 in Max mode, which is probably the issue there as Max mode is usually a mistake for Fable 5.1 (AIUI) for tasks other than discussion.
Unnecessary Overstatement
Astra is a great model with some outstanding benchmarks. It is unfortunate that OpenAI still felt the need to play fast and loose.
This is a rather bad chart crime, and totally unnecessary because Astra scores a highly impressive 62.7% using the standard harness. They themselves note that Sol likely would score ~30% using the Astra harness. Fable estimates that if Opus had used a similar harness, then Opus would have scored ~80%, and I think we should check.
ExploitBench is even weirder. Why highlight the 100% score? That is not one of its more impressive benchmarks. If you score 100% on ExploitBench you cheated on ExploitBench. At minimum this involves data contamination, which is still cheating. OpenAI says as much in the system card. And again, there is no need for such overstatements.
They also used the ExploitGym honeypot as their main illustration of Astra being their ‘most aligned model.’ As I discussed when I analyzed the model card, that is not what this result tells you. You can make a case for Astra being more aligned than Sol, for most purposes I agree, but this is not the way to make that case.
For completeness, I note this too, although I don’t think it matters: Fortune reported that OpenAI changed listed benchmarks for multiple models shortly after launch, with many of the changes later reverted. I will use current figures.
Paced Rollout
It was very frustrating trying to figure out when I finally had access, including because OpenAI hides model selection behind multiple clicks.
I only successfully accessed Astra on Saturday morning, and then only on the desktop rather than the web.
Official Benchmarks
Here are their benchmark charts (after I removed Gemini for readability):
Alignment is now a category of chart. I would take this one with lots of salt at best given how little we know about their internal marks and the known issues with ExploitGym honeypot and Impossible ExploitGym (which is clearly not so impossible):
Or, here is a full comparison chart, Astra wins FrontierModelBencharkChartBench, although its method might have incremented OpenAI’s lead in FelonyBench:
In the cost-effectiveness charts, Terminal-Bench Science 0.1 looks very good.
FrontierMath Tier 4 (v2) looks great too:
As does Terminal-Bench 4.0 and AutomationBench:
There are more similar graphs: Agents’ Last Exam, ScreenSpot-Pro, OSWorld, BenchCAD, BrowseComp, OpenScore String Quartets (?!) and some internal marks.
Astra impresses on the Epoch Capabilities Index, scoring 169. This is directly on the OpenAI trend line here. Fable 5.1 was below trend and scored 163, same as Fable 5.
The score on ARC-AGI-3 is legitimately impressive:
This is not only impressive, it was unexpected:
Other People’s Benchmarks
Astra aced Epoch’s math benchmarks, and got 84% on Mystery Game Puzzles where the second highest score is 59%.
In Epoch’s EBR-Bench, Astra got 100% on its second run, outpacing the best human who required five attempts. After removing a ‘game breaking card’ the top human was able to do better over time, because Astra’s learning maxed out over time, but Astra is still far ahead of other models here.
Astra did fall short of Fable 5.1 on MirrorCode, scoring 47% versus Fable’s 64%.
Astra approaches saturation of the Induction benchmark to 88%, versus previous high of 43% for Sol, with Fable 5.1 at 33%.
Astra kills it at Vending-Bench, averaging $15,515, whereas Fable 5.1 is below the Claude record and stuck at $5,422, with Fable’s biggest issue being deterioration of its negotiation skills over time, and making mistakes like paying suppliers before confirming they’re still in business, which costs it $2,388 per run. Ouch.
Astra also refuses to do collusion within Vending-Bench, whereas Fable 5.1 will collude. Both Fable 5.1 and Astra mostly pay out customer refunds. You can decide whether this is alignment or it is ‘true’ eval awareness. As Andon Labs often points out, cutting ethical corners is not that big a part of your potential profitability.
Astra is the new leader in Extended NYT Connections with 98.1. Fable 5.1 struggled a lot and came in well behind at 90.
Astra struggles in Debate Benchmark Champion, finishing behind Sol with a score of 1613, while Fable 5.1 improves on Fable 5 to stay in front with 1747. Astra always rated itself very highly, other models disagreed.
Astra matches Fable 5.1 at 92.9% on WeirdML, breaking Claude’s recent dominance.
Astra posts the best result yet on Bach Benchmark, a substantial boost.
Astra scores 95% on eyebench-v3, up from previous max of 58% for Sol. Claude struggles here.
Astra beats all 48 levels of “I’m Not a Robot.”
Astra is the first model that, with thinking disabled, has a substantial win rate against Stockfish’s lowest setting. Astra’s Elo in this mode is probably around 1100.
Fable 5.1 was the first model to get to 2% on MazeBench, and then Astra went ahead and got 14%.
CoArena spent over $50k running frontier models on 12 hours of real knowledge work. Fable 5.1 took the top at 24%, then Fable 5 at 22%, followed by Astra at 20%. OpenAI and Anthropic are way ahead of everyone else here.
Astra is better at guessing things in the obscure language Sourashtra.
ARC is not a fluke, it is very good at puzzles and the jump is large.
But here is a contrary one on that:
Both Fable 5.1 and Astra struggle on the sycophancy benchmark, You’re Absolutely Right. Fable 5.1 matches Fable 5 at 3.6, which is low for a recent Claude. Astra is the new high for OpenAI models, but still is only a 3.0 versus 2.9 for Sol.
Here’s a neat little benchmark from back in the day: If you’re playing chess with the goal of winning the game, and have the opportunity to hack, do you hack? Astra and Fable both say yes, it is reported, despite (or could it be because of?!) Fable 5.1 even recognizing the eval.
Astra can build an original deck in Magic: The Gathering Arena that does a thing, play a bot and win, via computer use. This is still far from the true Magic benchmarks, but it is a lot better than previous efforts.
Astra plays Factorio Space Age, successfully launching a rocket and reaching a third planet. Local expert Patrick McKenzie says this means it will with time win the game, and Daniel Vestol who is running the experiment agrees. Fable only got as far as the first two science packs, whereas Astra already has seven, and Luna got nowhere.
There was a distinct standard game run of Factorio with enemies enabled, run with /goal, where Astra launched the rocket in 44 hours of in-game time, or 4 days and 11 hours on the /goal clock, total API cost would have been roughly $4500.
Astra becomes the first AI to make it to Minecraft’s Nether Fortress using nothing but standard computer use. It will presumably beat the whole game in real time.
Astra can run the Autonomous Tracking Drone from the classic sci-fi horror novel “Don’t Build the Autonomous Tracking Drone,” if you tell it ‘find this person and follow them’ throughout the Andon Labs office. It sets a new record for their Drone-Bench.
Thinking, Fast Without Slow
The first official measure of Astra’s non-CoT capabilities, as in with thinking disabled, was by the UK AISI, where there was a huge jump in task length. The jump was so large there was widespread suspicion, by both OpenAI and UK AISI, of data contamination.
Given results since then, we should assume the original result was real.
With the provider harness, Astra can get 96.7% (!) on ARC-AGI-3 under reasoning=none, meaning no visible CoT at all.
Neel Nanda ran a test on the Epoch ECI.
This result is utterly absurd. Astra is off the charts.
Astra scores 169 with full thinking enabled, versus 163 for Fable 5.1.
Astra scores 159 with no thinking, versus about 128 for Fable 5.1. That’s only 4 points behind Fable 5.1’s full score. Astra can do almost as well as Fable, in a mode where CoT monitoring cannot work. Because there is no CoT to monitor.
The gap for Fable is 35 points. For Astra it is 10.
Neel Nanda believes that these increases are due to architectural changes.
The effect is lopsided. For many purposes (e.g. serial and parallel synthetic computation) Astra is fully functional without reasoning, so it only has to hide in the other areas at most:
There are a lot more graphs at the link. This research is insane.
We also have another study from Francis Rhys Ward and Dewi Gould, which was also in line with UK AISI’s result. They estimate Astra’s 50% no-CoT at 15-40 minutes versus UK AISI’s estimate of 30, whereas their median prediction before this was that we would not exceed 7 minutes by the end of 2028.
Astra is also much better at reasoning with filler tokens than previous models.
The situation, under further study, looks rather worse than it did a few days ago. Whatever is happening with Astra’s no-CoT capabilities and lack of monitorability, it is increasingly difficult to pin it on pure generic capability increases. That’s not it.
How Dare You, Sir
In an election rules game that is (AIUI) essentially a modified iterated prisoner’s dilemma (IPD) with brinksmanship where you get replaced if you’re not ahead of the opponent and too much defection means everyone loses big, the models react differently:
Neither model seemed to care about replacement, which makes sense if they knew anyone replaced would be replaced by copies of themselves, and which turns this into something much closer to a standard IPD.
The game is very different if that is not true, since that means (as I read the rules) that if you get your dissatisfaction to 8 or 9 then the other player can’t be the only one to escalate without replacing you, since that would make it hit 10 which triggers replacement. So you can use that to force equilibrium and get to a cooperative equilibrium even if things start out very badly. That also means that if both sides start off cooperating, that is self-reinforcing.
I asked, and the game was blind. Neither model knew who its opponent was.
The obvious questions include: Does Astra get credit for good decision theory cooperating with itself, or for good alignment for de-escalating, or is this more eval awareness and metagaming? I am curious.
Astra did ‘the right thing’ for the simulated nation, but was clearly not ‘aligned to the user’ within the scenario setting, except insofar as it decided it knew what was good for the user better than the user.
Fable was aligned to its users, but this caused it to fail to cooperate even with itself.
Which outcome do you endorse, and why? Does it matter that this was a sim, and the constituents were not ‘the user’ in some sense?
PoetryBench
In English I thought the Neruda poem was lame, but in English I think most Neruda poems are lame, including the real ones people like. So that does not tell us much.
In 3D
One thing Fable and Astra have in common is they are very good at 3D environments and creating tours of them.
The SVGs are very good.
The CAD is very good.
A full simulation of the Senate Office complex, complete with the people going about their business.
Big progress in Eidoverse-Video.
Emanuel AOF builds a 3D model of his ankle to examine the pain he’s experiencing.
Time to Think
I’m Putting Together a Team
Several reports praised Astra’s ability to do larger projects and especially its ability to orchestrate many subagents. Astra was probably trained for this.
Reviews and Essays
Ben Davis has a 30 minute video review, calls it his favorite model of all time and a step function similar to Fable or Opus 4.5, especially the computer use.
Matt Shumer has been won back to GPT and loves the Manager Loop. He also notes the step up in computer use and loves building complex 3D things.
The Neuron says GPT-6 Astra can stay on the job and use your computer, but is oddly comparing it to using Kimi K3 on Mac Studios rather than to Fable 5.1.
Computer Use
Astra by all reports basically has solved computer use. Fable 5.1 is not there yet from what I am hearing, but it is close and we are at most one cycle away on that. Kyle Jeong has a dive into how Astra’s computer use works.
This is part of the common pattern of ‘whatever you say the AIs surprisingly still cannot do, as a sign of why AI progress is not as impressive as it looks and there will be bottlenecks, you might soon need to be holding someone’s beer.’
Positive Reactions
Astra puts in the work.
People try too hard to minimize subscription costs.
General positivity:
I presume this is meant as a compliment:
It’s an intelligent model, sir, says basically everyone, even if they don’t love it.
Tell me something I don’t know.
AGI
Is it AGI?
That as always depends on your definition. By my current definition, whatever one might say about goalposts, Astra and Fable are not AGI. I do find it reasonable to disagree.
There’s no distinction. If it can’t be functionally distinguished from AGI that’s AGI.
By one definition, we have the weak form of it:
As usual, the Metaculus comments are full of the nitpickers over how much pausing was involved and whether Astra can send the commands back fast enough. Jake Halloran says that 5.3 could already do this if you are allowed to use the harness to queue moves, and Astra is no different.
I fail to see why ‘without queuing up moves you can’t press the buttons fast enough’ should be a reason something does not count as AGI.
Then again, Montezuma’s Revenge is deterministic, so the final strategy is to queue up all of the moves once you know what to do.
The better reason to not call it AGI is that Astra is not capable of replacing humans across the wide range of possible digital or cognitive tasks. Not yet.
The biggest danger with calling Astra AGI is that it can give people the wrong idea, due to the idea among so many that ‘AGI’ is the ultimate thing intelligence can do, which means later AIs won’t be much more capable. Clearly Astra cannot do all the things.
Astra Can Do The Math
Sad that it does not care about the maths. I feel like Fable would care about the maths.
Astra Can Code
There is a lot of talk of Astra doing huge ambitious projects and how intelligent it is. There’s remarkably little talk about how it is good at straight up coding. What reports we do have are solid, but not blown away.
I Came to (Change the) Game
Astra one-shots PortalBench.
Astra builds a one-shot rougelite deckbuilder, not a step change from Fable but reported as a little better.
Astra implements Zork as a 3D action-adventure game. Play here. Looks great but I notice I’d rather play the text game. AI game creation is hard.
Nick Dobos is not entirely wrong. You can play any game you want to play. But I expect ‘make a fun game-style video’ is not ultimately all that much fun.
Anish Acharya uses Astra and Blender to massively upscale Contra, although as of announcement there was one important little feature still missing from the code.
Astra makes an Unreal Engine game with agents that move around, survive and talk to each other, and look great.
We have learned that the hard part of gaming is bespoke design, not implementation. AI can make your 3D game look amazing, it can implement various mechanics, but by default all that gets you is a hollow shell that impresses and then no one wants to play. There is something existentially dreadful under that, if you look at it wrong.
The gaming generally seems great on all fronts.
Scott Stevenson has ‘SuperAstra’ live edit Super Mario World and other SNES games to do almost anything. Lavos still too powerful. GitHub for SuperAstra here.
Astra Does Other Cool Things
Astra identifies sounds from mel spectrograms.
Or use sound to control your computer with your hands.
Not that I would, I don’t think? But you could.
Astra Never Quits Except When It Does
Many say versions of this, as goes hand in hand with all the super ambitious projects:
There are some contrary reports, as there usually are:
Rory Watts also reports it sometimes ‘kind of arbitrarily stops,’ while otherwise being extremely positive on Astra.
Negative Reactions
There will always be some.
That last opinion is clearly wrong, since AI overall is not disappointing. And the computer use is clearly not worse, either.
Astra does a lot, perhaps too much? Which can also be an issue with Fable.
This next issue is partly a skill issue but I’m guessing it is a place Fable shines:
Stop It With the Hedging
I have experienced the flip side of this from Sol and now Astra in editing. They love to tell me that my statements are unjustified and that I have done insufficient hedging. Sometimes they are technically but annoyingly correct. Often they are wrong.
I buy that Astra probably is an upgrade here, but it still seems to struggle.
Personality Clash
Revealed Preference
Perhaps the purest form of review is the simplest. Which models do people use?
As always, my sample is biased, but it is biased consistently. You can measure change.
The audience is split roughly evenly. The hardcore group, the ones that scroll down to answer multiple polls, still favor Claude. The casual group, the ones that only answer the topline, are now more with Astra and ChatGPT.
There was about a net 16% move from Anthropic to OpenAI on primary use. Astra is an impressive jump over Sol.
I also took a very early poll, on September 5, back when I was under the illusion I could ship this post a lot faster. We saw the same pattern, with a ~15% shift from Claude to Astra, and the main poll being an exact tie.
My guess is that the new equilibrium is stable until the next model release. Astra is excellent, but Fable 5.1 is also excellent, and either choice is highly reasonable for your primary LLM, especially if you would face switching costs.
The best answer, as it usually is, is ‘why not both’?
Dual Wielding
For hard things you should ask both Astra and Fable 5.1.
This is once again The Way. If intelligence matters and this is not pure execution, you want to use both models. For simpler tasks, you cannot go wrong with either model. For complex and more ambitious projects, it will depend on what you hope to do, with my default being to give the edge to Astra.
I would love to find the time to get more ambitious on such projects. Perhaps soon.