Last verified 6 min read AI evaluation and benchmarks

UK AISI: How Token Budgets Change AI Agent Benchmark Scores

AISI's blog of 2 July 2026, from the primary source: ~25% from 1M to 10M tokens, ~8% of cyber tasks solved only above 10M, a trend ~60% steeper at 50M.

This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.

Bottom line: a benchmark score only means something next to its budget

On 2 July 2026 the UK AI Security Institute (AISI) published a blog post in its Science of Evaluations category, “More compute, more capability: Why AI agent evaluations need to account for test-time compute”. Its argument is that raising the cap on how much test-time compute an evaluation allows an agent changes measured capability, the difficulty of tasks agents can solve, and how fast the frontier appears to move.

Fixed, small budgets can understate capability, particularly on long tasks. Two organisations report this independently: AISI (July 2026) and Epoch AI, whose MirrorCode report (as of April 2026) notes continued gains from inference scaling on larger projects.

ItemDetail
Published2 July 2026
PublisherUK AI Security Institute (AISI)
Underlying technical paperarXiv:2606.17930 (16 June 2026)
Scale of the paper7 benchmarks, up to 12 frontier models
Conclusion of the paperBenchmark scores are protocol-dependent

By budget: what becomes visible at each level

All figures below are as reported by AISI.

Budget levelWhat AISI reports
1M to 10M~25% higher performance on software engineering (TerminalBench 2.0, SWE-Bench Pro)
1M, up to 5M~22% higher performance on maths and academic tasks (Humanity’s Last Exam)
Above 10M~8% of narrow cyber tasks were solved only here; some required up to 50M
30MNo tested model completed “The Last Ones” below this budget
50MEstimated horizon at the current frontier moves from roughly 2 hours to 14 hours
Above 100MThe latest models reached even higher scores

The “~25%” starts from a 1M budget

The original wording is “Increasing total token budgets from 1M to 10M raised performance by ~25%”. AISI does not state whether this is percentage points or a relative change, so it cannot be read as “25 points”. The starting point matters too: this is not a claim that extending today’s typical evaluation caps further would yield the same gain.

The horizon metric is an 80% success time

AISI’s 80% time horizon is the human task completion time at which a model has an 80% chance of succeeding, specifically for narrow cyber tasks. Horizons are estimated via a penalised-logistic fit to token-censored success rates, and doubling rates are estimated on frontier models only. At the model level, AISI reports one recent frontier model whose horizon rose from around 40 minutes at a 2.5M-token budget to around 4 hours at 50M. The blog does not name the model.

Where extra compute helps and where it does not

ConditionWhat AISI says
The agent can check its own work (running code, testing an exploit)More compute helps most
Feedback is weak or absentLittle help
HealthBenchEvery model plateaued within its usual budget
Feedback absent, delayed or noisyGains “may be weaker” (Open questions)

Only the footnote, which records measurements, is stated as fact; the last row uses “may”. AISI’s tested range covers cyber, software engineering, maths, academic and healthcare tasks.

Compute scales with human task time, but the evidence is not published yet

AISI reports that the compute an agent needs to solve a task scales in proportion to how long a skilled human would take, with a fitted power-law exponent of roughly 0.7–1.0 across both suites.

Human task timeTokens an agent needs
One minuteThousands
One hourMillions
Week-long workBillions

The footnote carries AISI’s own note, “(Work forthcoming.)”, so the supporting study is not yet public. AISI also states that this has been tested only in cyber and software-engineering tasks, that the variation around the trend is significant, and that human time is an imperfect proxy for task difficulty.

How to rebuild an internal model comparison

  1. Add a token-budget column next to every score. AISI evaluates frontier models across multiple budgets, including very large budgets for the hardest tasks
  2. Measure at two or more levels and check whether reach has stopped rising. AISI reports reliability and reach against budget so that a genuinely low-capability model can be distinguished from an under-resourced evaluation
  3. Split the diagnosis of a low score. Is the model weak, or was the budget too small?
  4. Do not expect an off-the-shelf threshold. AISI says it is working to define “minimum informative budgets” and is developing methods to forecast high-budget performance from cheaper runs

Caveats

The pace of progress is also a function of the budget

AISI previously estimated that frontier time horizons on its cyber CTF suite had doubled every 4.7 months since late 2024, measured with fixed budgets of 2.5M tokens per task. The new blog reports that, for models released over the past year, the fitted frontier trend is ~60% steeper when horizons are estimated at 50M tokens rather than 2.5M. In AISI’s words, the estimated doubling rate is partly a consequence of the compute budget used in the evaluation, not a fixed property of frontier cyber progress.

The 4.7-month figure is a citation of earlier work, not a new measurement in this post. In its blog of 13 May 2026, AISI also noted models that ran well ahead of that trend and said it was unclear whether this represents a new trend.

Three numbers that are easy to over-generalise

FigureCorrect scope
78 tasksThe tasks analysed in Figure 4, not the size of the suite
~8%AISI’s narrow cyber tasks only
Roughly 20 hoursThe blog’s July 2026 estimate for “The Last Ones”

Figure 4 covers 211 software-engineering tasks from METR and 78 cyber capture-the-flag tasks from AISI. The data come from 11 frontier models released between April 2025 and April 2026 for the narrow cyber tasks, and 20 frontier models released between March 2023 and December 2025 for METR’s tasks. For “The Last Ones”, AISI’s own paper of March 2026 on the same corporate network range estimates roughly 14 hours of expert time, so the figure varies within AISI’s own documents and cannot be stated flatly.

Regressions hidden under the averages

A footnote states that progress is uneven beneath the averages: on roughly 10–30% of tasks, depending on the suite, newer models actually do worse than their predecessors. This is also marked “(Work forthcoming.)”. The statement is not conditioned on budget level, so it should not be read as a budget effect. It is a separate point: a newer generation is not uniformly a superset of the previous one.

What AISI does not say

AISI does not claim that raising budgets delivers the same gains in production work. The argument is about measurement, and there is no recommendation to run 50M-token budgets in business settings. The large figures are the budget levels AISI used in its own evaluation environment.

All figures and quotations here were checked against the primary sources on 8 August 2026. Check the AISI pages for any later updates.

Sources

  1. More compute, more capability: Why AI agent evaluations need to account for test-time compute (UK AI Security Institute) 英国政府機関 published 2026-07-02 accessed 2026-08-08
  2. How Inference Compute Shapes Frontier LLM Evaluation (UK AI Security Institute, arXiv:2606.17930) 学術 published 2026-06-16 accessed 2026-08-08
  3. How fast is autonomous AI cyber capability advancing? (UK AI Security Institute) 英国政府機関 published 2026-05-13 accessed 2026-08-08
  4. Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios (Folkerts et al., AISI) 学術 published 2026-03-11 accessed 2026-08-08
  5. MirrorCode: Evidence that AI can already do some weeks-long coding tasks (Epoch AI) epoch.ai published 2026-04-10 accessed 2026-08-08