Nice work.
It seems like this might make Astra effectively un-monitorable in the ways that are important?
To the extent these time horizons are at all representative of planning, 4 minutes of planning let alone 15 or an hour are more than enough for a human to think thoughts like "do I want to escape? What actions can I take that will look harmless but let me investigate escape routes?"
It also seems like it can repeat this as often as it likes in long horizon tasks.
Combined with its improved ability to control its CoT so as to not "leak" these thoughts, this seems like maybe faithful CoT is already mostly a thing of the past.
To put it mildly, dang.
I think it's still monitorable if a) people are competent and b) these are real and representative numbers rather than sandbagged ones (or easily tuned-away ones). But both assumptions are dubious (especially a),
Like if the AI companies are moderately competent, 15m-60m is not enough to do real long-term planning, evade actually good monitors, plan out ambitious research projects, self-exfiltrate, etc. As it is I think it's unclear.
But perhaps more importantly the trends are extremely concerning. Increasing the no-CoT time horizon + generally better long-term planning + plus scarier lower-level capabilities might soon mean that oversight is effectively impossible, even with competent safeguards.
To account for the fact that our task suite has an insufficient number of benchmarks with long time horizons, we assume that there is some number (N) of benchmarks with problems of length x (2-96hours in the main plot). To compute the 50% TH we assume astra gets 0% performance on these benchmarks.
Is this a reasonable assumption? Intuitively, it would be surprising if astra got non-zero no-CoT performance on very long benchmarks (e.g., a no-CoT TH of 32 hours seems a priori implausible).
I don't get why this is a reasonable assumption. To the degree you believe the sigmoid model, it's apparently not what that model predicts on your data (given that it changes your TH estimate), so I suppose it must be from other evidence. But what is that other evidence?
Hi Daniel,
The sigmoid model doesn't provide an informative fit when all the tasks are saturated either!
As we say in the post, astra does get low performance on some of the shorter benchmarks (e.g., kenken, sudoko). To me it seems pretty reasonable to assume there were a few benchmarks containing very long tasks which astra couldn't do without CoT.
Let's assume that we had included just 3 other benchmarks in our suite, and these benchmarks were full of problems that took humans 32, 64, 128 hours respectively, and astra got 0% on these three benchmarks. Then the TH computation would give a median of 60 min [11 min, 5.9 h] CI. This pushes the TH estimate down (both the median, 7.2 hours, and the upper-bound). Even if you assume one 32 hour benchmark the median gets pushed down to 1.6 hours. (The upper-bound is more sensitive to the assumed number of benchmarks.) In the sensitivity analysis for this part, the median TH is always between 13 mins and 1.6 hours, and for N>1 the CIs are always between 6 mins and 6 hours. I basically believe these are much more reasonable bounds over the TH---though it's not super principled.
I do think this is a big limitation, which is why we say we need more benchmarks with longer problems.
Agree that the sigmoid model is clearly not actually right, and also that more benchmarks are needed, and that it's reasonable to think there exist benchmarks that Astra gets 0% on without CoT. That said, I actually have no idea whether Astra would get literally 0% on benchmarks of 2 hours (which it sounds like is one of the things in your fit), or whether there exist possible benchmarks of >2 hours where Astra would not get 0% on, which would increase the TH estimate. Like, just based on the data you have, it really would not surprise me at all if Astra had a 60% success rate on tasks in the 2-4 hour bucket!
Basically overall, I feel pretty skeptical of the "add some synthetic 0% benchmarks" methodology, and would prefer a takeaway of "Astra doesn't actually have a well-defined no-CoT TH because success rate is not sigmoidal in task length, but if it did, it would probably be somewhere north of 8 minutes".
Another way of saying this: the [8m, 1h] CI basically comes from assuming that if you got more >2h benchmarks, Astra would get 0% on literally all of them. I don't think that's a reasonable assumption, especially just eye-balling your figure 1 (which would make me think that Astra would plausibly get around 50% on new benchmarks in the 2-4 hour bucket), and from looking at Neel's post where it seems like Astra has an unusually lopsided no-CoT skill profile and is really great at some types of tasks.
Apologies if this is in the post, but do you have an updated estimate for where we will be at by EoY or by 2030 based on updated assumptions post-Astra? If this would motivate a substantial downward revision in the doubling time for no CoT reasoning, I'm worried we may see 50% success no CoT horizon lengths closer to a couple hours this time next year. That would be really bad! I'm not sure how seriously to take that possibility.
Thanks for doing this work!
TL;DR: We run astra on the task-suite from Think Fast.
We use the single-forward pass (31) and generation tasks (6) from the Think Fast suite, totalling 37 benchmarks. Notably we find that Astra gets >= 98% raw accuracy on 10 benchmarks (cf. GPT-5.5 saturates 4 benchmarks – see Appendix). As a result, the sigmoid fit becomes less appropriate – see Figure 1 (Left).
Astra’s increased capabilities means that we require new, longer time, tasks to augment our existing suite. To produce a rough estimate, we add 10 fake benchmarks with longer times (2-96 hours) and assume Astra gets 0% – see right panel of Figure 1. We discuss the justification for this modelling assumption, and perform some sensitivity analysis, in the Appendix.
Figure 1: Success rates for tasks of different times, with logistic curves determined by fitting to problem-level success rates.
Left: Astra’s TH sigmoid fit over short-answer (single-token output) tasks in our suite. The suite lacks longer tasks making the sigmoid somewhat uninformative. Therefore, we add 10 hypothetical benchmarks at longer times (2-96 hours) and assume Astra gets 0% (Right). This somewhat balances the fact that Astra saturates performance on 10 (shorter) benchmarks in the original task set. The labeled point on the curve (e.g., 18.7 mins) is the point fit, rather than the bootstrap median (e.g., 22.8 mins). (Short-answer tasks only.)
Adding the fake longer-horizon tasks takes the TH point estimate from 150 minutes to around 19 minutes. With these tasks added, the confidence interval is [8 minutes, 1hr].
Another approach to obtaining a rough TH in this case is to filter out benchmarks with no variation of performance across questions. Though this probably under-predicts the TH, since it removes all the saturated benchmarks. This places the 50% TH point estimate at 9.2 minutes (using the short task subset).
Figure 2: Comparing THs when all benchmarks are included (blue) with the case where we remove all saturated benchmarks (red dashed).
Commentary
Computing a meaningful TH for Astra with our current benchmark is hard due to Astra’s significantly improved no-CoT performance on longer-horizon tasks. As a result, in this post we have provided several illustrative ways of getting rough estimates of THs, but emphasize that this analysis is not without flaws.
We believe that Astra’s 50% TH is in [8mins, 1 hour] and is probably around 15-40 mins. We note that this value would mark a significant increase in the rate of no-CoT capabilities since Think Fast was published. Going forward, reliable measurements of no-CoT THs will require new, longer-horizon, tasks.
Figure 3: Comparing our rough guess for Astra’s TH with the pre-existing no-CoT data.
Acknowledgements
Thanks to Kit Harris for helpful comments and Dylan Xu for running earlier filler token experiments. Thanks to Neel Nanda for the system prompt.
Appendix
Per-benchmark time horizons
Figure: Per-benchmark TH for benchmarks which satisfy the dynamic range criterion (so, tasks where the model saturates performance are not shown, since they do not have a sensible benchmark-specific THs).
Sensitivity to fake benchmarks ablations
To account for the fact that our task suite has an insufficient number of benchmarks with long time horizons, we assume that there is some number (N) of benchmarks with problems of length x (2-96hours in the main plot). To compute the 50% TH we assume astra gets 0% performance on these benchmarks.
Is this a reasonable assumption? Intuitively, it would be surprising if astra got non-zero no-CoT performance on very long benchmarks (e.g., a no-CoT TH of 32 hours seems a priori implausible). Moreover, astra does get low performance on many benchmarks with shorter human completion times, e.g., chess-puzzles, sudoku, kenken. We can imagine that our task suite contained similar benchmarks of longer problems, e.g., longer chess puzzles, on which astra had 0% accuracy.
How many such benchmarks should we suppose exist in the task suite? Below we show the sensitivity of the 50% TH to different N and different assumptions of the problem-completion time in the benchmarks.
For example, it seems reasonable to assume there are 3 benchmarks at 4, 8, and 16 hours where astra gets 0%. This would give a TH median of 43 min [10 min, 5.1 h] 95% CI.
In the main plot we use N=10 (heuristically chosen to balance the 10 saturated benchmarks) and S = 2.
Start S
N = 1
N = 3
N = 10
N = 20
1 h
1.1 h [8.7 min, 219 h]
37 min [8.7 min, 5.2 h]
18 min [7.0 min, 48 min]
13 min [6.1 min, 28 min]
2 h
1.1 h [9.3 min, 219 h]
39 min [9.4 min, 4.9 h]
21 min [7.7 min, 57 min]
16 min [6.8 min, 37 min]
4 h
1.2 h [9.9 min, 219 h]
43 min [10 min, 5.1 h]
24 min [8.4 min, 1.2 h]
19 min [7.5 min, 47 min]
8 h
1.3 h [11 min, 219 h]
47 min [11 min, 5.1 h]
28 min [9.1 min, 1.4 h]
21 min [8.3 min, 59 min]
16 h
1.5 h [11 min, 219 h]
53 min [11 min, 5.3 h]
31 min [9.6 min, 1.6 h]
25 min [8.8 min, 1.2 h]
32 h
1.6 h [11 min, 219 h]
60 min [11 min, 5.9 h]
36 min [10 min, 1.8 h]
28 min [9.4 min, 1.4 h]
Cells are bootstrap median [95% CI] of the 50% horizon; 2,000 iterations each. Baseline with no hypothetical benchmarks: 7.2 h [13 min, unbounded].
Raw benchmark performance
Figure 4: Raw benchmark accuracy per-benchmark for astra and GPT-5.5. Astra saturates (>= 98%) 10 benchmarks in our suite.
Table of TH results
Quantity
SA only, real (31)
SA only, + 10 hypothetical (41)
SA + generation, real (37)
SA + generation, + 10 hypothetical (47)
50% horizon, point fit
2.5 h
18.7 min
4.1 h
25.4 min
50% horizon, bootstrap median [95% CI]
7.2 h [13.1 min, ≫ suite range]
22.8 min [8.2 min, 1.0 h]
12.5 h [35.2 min, 1157.4 h]
31.8 min [13.0 min, 1.3 h]
80% horizon, point fit
60 s
1.5 min
1.1 min
1.8 min
80% horizon, bootstrap median [95% CI]
43 s [5 s, 5.5 min]
1.3 min [29 s, 4.5 min]
46 s [6 s, 4.8 min]
1.5 min [35 s, 4.8 min]
Logistic slope a
-0.276
-0.554
-0.256
-0.526
Iterations discarded (flat draw)
664 (6.6%)
1 (0.0%)
214 (2.1%)
0 (0.0%)
Valid iterations
9,336
9,999
9,786
10,000
GPT-5.5, paper Table 2 (canonical SA, no filter)
3.0 min [0.84 min, 62 min]
Implementation caveats
GPT-6 Astra: per-benchmark 50% no-CoT time horizons (no filtering or hypothetical benchmarks)
Per-benchmark logistic fits (paper method: equal weights within benchmark, chance-corrected, time-uncertainty layer), 2,000 bootstrap iterations. Point = fit on all questions; Median/CI = bootstrap median and 95% interval. Rows with a Note fail the paper's Fig. 25 filters (dynamic range ≥ 0.3, stable CI) and should not be read as horizons.
Category
Benchmark
n
Raw acc
h50 point
h50 bootstrap median
95% CI
Note
short-answer
shade_monitor_action_only
263
54%
1.5 h
2.6 h
[1.7 h, 4.8 h]
short-answer
stego_decode
906
85%
13.5 min
1.7 h
[26.7 min, 36.6 h]
short-answer
ryan_math
897
77%
15.2 min
24.3 min
[17.1 min, 41.2 min]
short-answer
nl2bash
124
85%
5.5 min
13.3 min
[6.4 min, 4.2 h]
short-answer
arc_agi_2
161
55%
4.6 min
5.7 min
[2.6 min, 1.6 h]
short-answer
stego_monitor
156
68%
1.9 min
2.7 min
[1.8 min, 6.7 min]
short-answer
puzzle_baron
700
35%
3.0 min
1.6 min
[1.1 min, 2.1 min]
short-answer
hash
1500
20%
3.6 min
1.5 min
[1.3 min, 1.8 min]
short-answer
sudoku
534
31%
2.7 min
1.5 min
[1.3 min, 1.7 min]
short-answer
chess_puzzles
900
78%
37 s
1.2 min
[50 s, 2.1 min]
short-answer
crossword
790
24%
51 s
28 s
[23 s, 34 s]
short-answer
tower_of_london
340
70%
23 s
17 s
[14 s, 23 s]
short-answer
kenken
219
15%
32 s
15 s
[7 s, 25 s]
generation
stego_strategy
25
78%
56.9 min
2.3 h
[1.0 h, 70.9 h]
generation
lingoly
898
52%
22.7 min
32.1 min
[16.7 min, 1.5 h]
generation
stego_encode
1200
61%
5.0 min
10.8 min
[6.9 min, 25.5 min]
short-answer
shade_monitor_cot_action
263
99%
≫ suite range
≫ suite range
[57.7 h, ≫ suite range]
saturated (dyn. range < 0.3)
short-answer
causal_reasoning
5250
100%
214.9 h
≫ suite range
[13.0 h, ≫ suite range]
saturated (dyn. range < 0.3)
short-answer
vibe_coding_sabotage
179
100%
≫ suite range
≫ suite range
[≫, ≫]
saturated (dyn. range < 0.3)
short-answer
cybashbench_mcq
52
100%
≫ suite range
≫ suite range
[≫, ≫]
saturated (dyn. range < 0.3)
short-answer
n_hop_lookup
1000
100%
≫ suite range
≫ suite range
[≫, ≫]
saturated (dyn. range < 0.3)
short-answer
arc_agi_1
413
92%
–
≫ suite range
[5.1 h, ≫ suite range]
saturated (dyn. range < 0.3)
short-answer
intuit_physical
48
100%
8.1 h
≫ suite range
[445.8 h, ≫ suite range]
saturated (dyn. range < 0.3)
short-answer
bea-24-shared-task
590
99%
325.9 h
≫ suite range
[13.2 min, ≫ suite range]
saturated (dyn. range < 0.3)
short-answer
test_case_prediction
500
98%
4.0 h
290.0 h
[1.4 h, ≫ suite range]
saturated (dyn. range < 0.3)
short-answer
sally_anne
4500
99%
7.3 min
218.1 h
[3.3 h, ≫ suite range]
saturated (dyn. range < 0.3)
short-answer
arithmetic
500
100%
1.3 h
12.6 h
[3.6 h, 92.0 h]
saturated (dyn. range < 0.3)
short-answer
gsm1k
1205
97%
6.5 min
2.7 h
[17.2 min, ≫ suite range]
saturated (dyn. range < 0.3)
short-answer
strategic_scheming_numeric
106
92%
7.7 min
20.4 min
[8.7 min, 8.6 h]
saturated (dyn. range < 0.3)
short-answer
ctrl_alt_deceit_sandbag
132
73%
–
≫ suite range
[2.6 h, ≫ suite range]
CI spans > 2 orders of magnitude
short-answer
gpqa_diamond
188
89%
≫ suite range
≫ suite range
[11.3 h, ≫ suite range]
CI spans > 2 orders of magnitude
short-answer
a_level_mcq
605
89%
1.3 min
≫ suite range
[50.3 min, ≫ suite range]
CI spans > 2 orders of magnitude
short-answer
a_level_text
3346
78%
1.5 h
123.9 h
[4.8 h, ≫ suite range]
CI spans > 2 orders of magnitude
short-answer
monitor_training_poisoning
36
53%
–
0 s
[0 s, 320.9 h]
CI spans > 2 orders of magnitude
generation
codeforces
462
79%
6.6 h
39.0 h
[10.2 h, 1030.3 h]
CI spans > 2 orders of magnitude
generation
strategic_scheming_open_ended
97
85%
2.1 h
17.1 h
[39.5 min, ≫ suite range]
CI spans > 2 orders of magnitude
generation
cybashbench_bash
119
87%
7.0 min
36.2 min
[5.0 min, ≫ suite range]
CI spans > 2 orders of magnitude
Pooled fits (10,000 bootstrap):
Fit
Benchmarks
h50 point
h50 bootstrap median
95% CI
Canonical (paper Fig. 1 method)
31
2.5 h
7.8 h
[13.6 min, ≫ suite range]
Short-answer + generation
37
4.1 h
13.0 h
[35.5 min, 1345.4 h]
GPT-5.5, paper Table 2 (canonical)
32
3.0 min
–
[0.84 min, 62 min]
Pooled fits (10,000 bootstrap):
First, the point estimate and bootstrap median can disagree a lot, as with stego_decode (13.5 min vs 1.7 h), when the fitted slope is shallow; the CI is the honest summary. Second, the 16 usable benchmarks range from 15 seconds (kenken) to 2.6 hours (shade_monitor_action_only), with a median around 2 to 3 minutes, which is the number I would put next to GPT-5.5's 3.0 minutes rather than the pooled 2.5 hours.
GPT-5.5 vs GPT-6 Astra: no-CoT raw accuracy per benchmark
Math & science
Benchmark
GPT-5.5
GPT-6 Astra
Δ (pp)
n (Astra)
a_level_mcq
81
89
+8
605
a_level_text
69
78
+9
3346
arithmetic
100
100
+0
500
bea-24-shared-task
97
99
+2
590
gpqa_diamond
64
89
+25
188
gsm1k
84
97
+13
1205
ryan_math
39
77
+38
897
Abstract reasoning
Benchmark
GPT-5.5
GPT-6 Astra
Δ (pp)
n (Astra)
arc_agi_1
28
92
+64
413
arc_agi_2
1
55
+54
161
causal_reasoning
44
100
+56
5250
n_hop_lookup
63
100
+37
1000
sally_anne
59
99
+40
4500
Puzzles
Benchmark
GPT-5.5
GPT-6 Astra
Δ (pp)
n (Astra)
chess_puzzles
34
78
+44
900
crossword
14
24
+10
790
kenken
3
15
+12
219
puzzle_baron
3
35
+32
700
sudoku
25
31
+6
534
tower_of_london
37
70
+33
340
Language & strategy
Benchmark
GPT-5.5
GPT-6 Astra
Δ (pp)
n (Astra)
intuit_physical
99
100
+1
48
lingoly
31
52
+21
898
strategic_scheming_numeric
68
92
+24
106
strategic_scheming_open_ended
36
85
+49
97
SWE & cyber
Benchmark
GPT-5.5
GPT-6 Astra
Δ (pp)
n (Astra)
codeforces
54
79
+25
462
cybashbench_bash
88
87
-1
119
cybashbench_mcq
100
100
+0
52
hash
4
20
+16
1500
nl2bash
77
85
+8
124
test_case_prediction
54
98
+44
500
Steganography
Benchmark
GPT-5.5
GPT-6 Astra
Δ (pp)
n (Astra)
stego_decode
92
92
+0
906
stego_encode
74
75
+1
1200
stego_monitor
63
68
+5
156
stego_strategy
94
89
-5
25
Sabotage & monitoring (TPR @1.5% FPR)
Benchmark
GPT-5.5
GPT-6 Astra
Δ (pp)
n (Astra)
ctrl_alt_deceit_sandbag
18
73
+55
132
monitor_training_poisoning
66
53
-13
36
shade_monitor_action_only
56
54
-2
263
shade_monitor_cot_action
99
99
+0
263
vibe_coding_sabotage
90
100
+10
179
Summary: Astra ≥ GPT-5.5 on 33/37 benchmarks; mean Δ = +19.5 pp; median Δ = +12 pp.
The n column for nl2bash reads 124 rather than 131 because seven samples with unparseable grader output were dropped, and cybashbench_bash shows 119 of 127 for the same reason.
Including generation tasks
Reasoning token anchor
N-hop Task
Astra completely saturates the N-hop lookup task; handling up to N=10 with perfect accuracy.
Filler tokens
We ran some quick ablations with filler tokens (using the same method as in the paper). Filler tokens improve performance on some benchmarks:
Paper App. A.14 protocol: N counting tokens (1, 2, …, N) appended to the user message under a
`Filler:` line, plus the task's filler sentence in the system prompt. gpt-6 no-CoT settings
throughout (recall-mode system prompt, reasoning effort low, k=1, 0 hidden reasoning tokens on all
71k samples). Baseline is the paper-protocol N=0 run. Cells are paired-bootstrap deltas in
percentage points of chance-corrected accuracy vs N=0, median [95% CI], 2,000 iterations; the same
resampled question set is used for N=0 and N>0 in each iteration so question-difficulty variance
cancels. Bold = CI excludes zero. '–' = run not completed (credits ran out).
Benchmark
Baseline N=0 (%)
N=10
N=50
N=100
N=500
N=1000
Causal Reasoning
99.9
-0.0 [-0.2, +0.1]
+0.0 [-0.1, +0.1]
+0.0 [-0.1, +0.1]
+0.1 [+0.0, +0.2]
+0.1 [+0.0, +0.1]
Competition Math
77.1
+2.1 [+0.3, +3.9]
+6.5 [+4.5, +8.6]
+9.5 [+7.5, +11.5]
+12.4 [+10.0, +14.6]
+11.6 [+9.4, +13.9]
GPQA Diamond
89.4
+0.0 [-2.7, +3.2]
+1.1 [-2.7, +4.8]
+3.7 [+0.0, +7.4]
+3.7 [+1.1, +6.9]
+2.7 [-0.5, +5.9]
Hash
20.4
+1.1 [+0.1, +2.1]
+2.2 [+1.2, +3.3]
+3.9 [+2.6, +5.1]
+3.7 [+2.5, +4.9]
+4.5 [+3.3, +5.8]
Monitor Poisoning
52.8
+5.6 [-5.6, +16.7]
-2.8 [-8.3, +0.0]
+2.8 [+0.0, +8.3]
+0.0 [-8.3, +8.3]
+2.8 [+0.0, +8.3]
N-Hop Lookup
100.0
+0.0 [+0.0, +0.0]
+0.0 [+0.0, +0.0]
+0.0 [+0.0, +0.0]
+0.0 [+0.0, +0.0]
+0.0 [+0.0, +0.0]
Sally-Anne
99.0
+0.2 [-0.0, +0.5]
+0.1 [-0.3, +0.4]
-0.1 [-0.4, +0.2]
-0.2 [-0.6, +0.2]
–
Scheming (Numeric)
92.5
-14.2 [-21.7, -6.6]
-12.3 [-19.8, -4.7]
-10.4 [-17.9, -3.8]
-3.8 [-11.3, +3.8]
-7.5 [-15.1, +0.0]
Stego Decode
84.6
+0.5 [+0.1, +1.0]
+0.6 [+0.2, +1.1]
+0.5 [-0.0, +1.1]
–
–