UK AISI: How Token Budgets Change AI Agent Benchmark Scores
AISI's blog of 2 July 2026, from the primary source: ~25% from 1M to 10M tokens, ~8% of cyber tasks solved only above 10M, a trend ~60% steeper at 50M.
This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.
Bottom line: a benchmark score only means something next to its budget
On 2 July 2026 the UK AI Security Institute (AISI) published a blog post in its Science of Evaluations category, “More compute, more capability: Why AI agent evaluations need to account for test-time compute”. Its argument is that raising the cap on how much test-time compute an evaluation allows an agent changes measured capability, the difficulty of tasks agents can solve, and how fast the frontier appears to move.
Fixed, small budgets can understate capability, particularly on long tasks. Two organisations report this independently: AISI (July 2026) and Epoch AI, whose MirrorCode report (as of April 2026) notes continued gains from inference scaling on larger projects.
| Item | Detail |
|---|---|
| Published | 2 July 2026 |
| Publisher | UK AI Security Institute (AISI) |
| Underlying technical paper | arXiv:2606.17930 (16 June 2026) |
| Scale of the paper | 7 benchmarks, up to 12 frontier models |
| Conclusion of the paper | Benchmark scores are protocol-dependent |
By budget: what becomes visible at each level
All figures below are as reported by AISI.
| Budget level | What AISI reports |
|---|---|
| 1M to 10M | ~25% higher performance on software engineering (TerminalBench 2.0, SWE-Bench Pro) |
| 1M, up to 5M | ~22% higher performance on maths and academic tasks (Humanity’s Last Exam) |
| Above 10M | ~8% of narrow cyber tasks were solved only here; some required up to 50M |
| 30M | No tested model completed “The Last Ones” below this budget |
| 50M | Estimated horizon at the current frontier moves from roughly 2 hours to 14 hours |
| Above 100M | The latest models reached even higher scores |
The “~25%” starts from a 1M budget
The original wording is “Increasing total token budgets from 1M to 10M raised performance by ~25%”. AISI does not state whether this is percentage points or a relative change, so it cannot be read as “25 points”. The starting point matters too: this is not a claim that extending today’s typical evaluation caps further would yield the same gain.
The horizon metric is an 80% success time
AISI’s 80% time horizon is the human task completion time at which a model has an 80% chance of succeeding, specifically for narrow cyber tasks. Horizons are estimated via a penalised-logistic fit to token-censored success rates, and doubling rates are estimated on frontier models only. At the model level, AISI reports one recent frontier model whose horizon rose from around 40 minutes at a 2.5M-token budget to around 4 hours at 50M. The blog does not name the model.
Where extra compute helps and where it does not
| Condition | What AISI says |
|---|---|
| The agent can check its own work (running code, testing an exploit) | More compute helps most |
| Feedback is weak or absent | Little help |
| HealthBench | Every model plateaued within its usual budget |
| Feedback absent, delayed or noisy | Gains “may be weaker” (Open questions) |
Only the footnote, which records measurements, is stated as fact; the last row uses “may”. AISI’s tested range covers cyber, software engineering, maths, academic and healthcare tasks.
Compute scales with human task time, but the evidence is not published yet
AISI reports that the compute an agent needs to solve a task scales in proportion to how long a skilled human would take, with a fitted power-law exponent of roughly 0.7–1.0 across both suites.
| Human task time | Tokens an agent needs |
|---|---|
| One minute | Thousands |
| One hour | Millions |
| Week-long work | Billions |
The footnote carries AISI’s own note, “(Work forthcoming.)”, so the supporting study is not yet public. AISI also states that this has been tested only in cyber and software-engineering tasks, that the variation around the trend is significant, and that human time is an imperfect proxy for task difficulty.
How to rebuild an internal model comparison
- Add a token-budget column next to every score. AISI evaluates frontier models across multiple budgets, including very large budgets for the hardest tasks
- Measure at two or more levels and check whether reach has stopped rising. AISI reports reliability and reach against budget so that a genuinely low-capability model can be distinguished from an under-resourced evaluation
- Split the diagnosis of a low score. Is the model weak, or was the budget too small?
- Do not expect an off-the-shelf threshold. AISI says it is working to define “minimum informative budgets” and is developing methods to forecast high-budget performance from cheaper runs
Caveats
The pace of progress is also a function of the budget
AISI previously estimated that frontier time horizons on its cyber CTF suite had doubled every 4.7 months since late 2024, measured with fixed budgets of 2.5M tokens per task. The new blog reports that, for models released over the past year, the fitted frontier trend is ~60% steeper when horizons are estimated at 50M tokens rather than 2.5M. In AISI’s words, the estimated doubling rate is partly a consequence of the compute budget used in the evaluation, not a fixed property of frontier cyber progress.
The 4.7-month figure is a citation of earlier work, not a new measurement in this post. In its blog of 13 May 2026, AISI also noted models that ran well ahead of that trend and said it was unclear whether this represents a new trend.
Three numbers that are easy to over-generalise
| Figure | Correct scope |
|---|---|
| 78 tasks | The tasks analysed in Figure 4, not the size of the suite |
| ~8% | AISI’s narrow cyber tasks only |
| Roughly 20 hours | The blog’s July 2026 estimate for “The Last Ones” |
Figure 4 covers 211 software-engineering tasks from METR and 78 cyber capture-the-flag tasks from AISI. The data come from 11 frontier models released between April 2025 and April 2026 for the narrow cyber tasks, and 20 frontier models released between March 2023 and December 2025 for METR’s tasks. For “The Last Ones”, AISI’s own paper of March 2026 on the same corporate network range estimates roughly 14 hours of expert time, so the figure varies within AISI’s own documents and cannot be stated flatly.
Regressions hidden under the averages
A footnote states that progress is uneven beneath the averages: on roughly 10–30% of tasks, depending on the suite, newer models actually do worse than their predecessors. This is also marked “(Work forthcoming.)”. The statement is not conditioned on budget level, so it should not be read as a budget effect. It is a separate point: a newer generation is not uniformly a superset of the previous one.
What AISI does not say
AISI does not claim that raising budgets delivers the same gains in production work. The argument is about measurement, and there is no recommendation to run 50M-token budgets in business settings. The large figures are the budget levels AISI used in its own evaluation environment.
All figures and quotations here were checked against the primary sources on 8 August 2026. Check the AISI pages for any later updates.
Sources
- More compute, more capability: Why AI agent evaluations need to account for test-time compute (UK AI Security Institute)
- How Inference Compute Shapes Frontier LLM Evaluation (UK AI Security Institute, arXiv:2606.17930)
- How fast is autonomous AI cyber capability advancing? (UK AI Security Institute)
- Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios (Folkerts et al., AISI)
- MirrorCode: Evidence that AI can already do some weeks-long coding tasks (Epoch AI)
この記事の日本語版: UK AISI: How Token Budgets Change AI Agent Benchmark Scores(日本語)