Last verified 5 min read LLM evaluation and evaluation cost

optstop 57-97% savings: the denominator behind the number

optstop reports 57-97% savings on planned trials. The range comes from nine shadow-mode settings, and the README states where its intervals are calibrated.

This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.

Conclusion: the 57-97% figure is a counterfactual, not a live result

On 27 August 2026 the UK AI Security Institute (formerly the AI Safety Institute) published a blog post introducing optstop, an open source tool that stops LLM evaluation sampling on statistical criteria. The headline range is 57% to 97% of planned trials. Three things define what that range is.

  1. According to the accompanying paper (arXiv:2608.14425, by Toby D. Pilditch), the range was measured in shadow mode: every planned trial is still executed, while the run internally records the point at which the stopping criteria would have been met
  2. The denominator is “planned trials across nine validation settings”, tied to an illustrative 200-item, 10-epoch evaluation design
  3. The README, not the blog, states where the credible intervals are well calibrated: 50 or more items per grouping, with performance in the 0.2-0.8 range

Every number below comes from the UK AISI and the paper’s author. As of 1 September 2026 no independent third-party replication or verification has been found.

The three primary documents word the savings differently

DocumentDateHow the savings are statedCaveat carried in the same document
AISI blog2026-08-27Two sentences: “under every condition we set … 57% and 97% of planned runs” and “under every condition … 57% and 97% of planned trials”With a leaner design and fewer trials, the number of trials saved may be lower
Paper abstract2026-08-14“removes 57%-97% of planned trials across nine validation settings”“with the magnitude of savings depending on evaluation design”
README2026-08-19No savings figure; calibration conditions insteadFor small samples or near-boundary performance, “CIs may undercover”

The blog’s two sentences are not identical. The first carries the qualifier “we set” and refers to planned runs; the second drops the qualifier and refers to planned trials.

Measurement conditions behind the range

The paper reports a 3x3 matrix experiment crossing three inference pathways (binary, ordinal, continuous) with three performance levels.

ItemValue
Validation settings (cells)9
Items per cell200
Epochs per item10
Planned trials per cell2,000
Precision threshold δ0.05
Credible interval level97%
Mean efficiency gain81.1%
Mean absolute score deviation0.006

One cell differs: mid-binary uses the full GPQA Diamond set, so it has 198 items and 1,980 planned trials. The 0.006 deviation is a mean absolute value on a normalised [0,1] scale, comparing full-run and truncated estimates. All nine cells triggered early stopping.

Stopping rules, and what happens when neither fires

Per the blog, sampling stops when either rule is satisfied.

  • Precision: the credible interval is narrow enough
  • Stabilisation: the interval has stopped changing, so the data has nothing more to tell us

If neither rule is met, the grouping simply runs to its full planned budget. A genuinely noisy estimate is not cut short; it is sampled to the end. A conservatism mechanism also demands more data when estimated success falls below 1% — that threshold is the default in version 0.4.0, whose CHANGELOG entry is dated 18 March 2026.

Adoption path described by the blog

  1. Post-hoc: run it over a completed evaluation and confirm the stopped estimates match the full-run estimates
  2. Shadow mode: run the evaluation as normal while optstop reports the saving it would have made, with the full dataset preserved
  3. Live: switch on live stopping once the shadow-mode numbers are convincing

The package ships an OptimalStoppingManager implementing the EarlyStopping protocol of inspect_ai, the UK AISI evaluation framework; the shadow_mode parameter defaults to False. The blog adds one practical requirement: tasks must be presented in a randomised order so that early estimates are not biased by queue ordering, and it states that Inspect supports this.

Caveats

Intervals are called well-calibrated only within a stated range

ConditionWhat the README says
50+ items per grouping, performance 0.2-0.8Credible intervals are well calibrated
Smaller samples, or performance close to 0% or 100%Intervals may undercover, due to hierarchical shrinkage
Stress test: 18-30 items per grouping with many boundary performersObserved coverage can fall substantially below nominal

Only the third row is stated as observed; the second is a possibility. The README frames this as a structural property of Bayesian hierarchical models rather than an implementation bug, and notes it affects live and post-hoc modes equally.

27 August 2026 is the date it was written about, not released

EventDate
Repository created (GitHub API metadata)2025-12-12
Latest CHANGELOG entry, 0.4.02026-03-18
README commit “Initial public release”2026-06-12
README last updated2026-08-19
Introduced on the AISI blog2026-08-27

The current version is 0.4.0 and the licence is MIT, both as retrieved on 1 September 2026.

Sources

  1. Optimal stopping: spending evaluation compute where it counts (AISI Blog) 英国政府機関 published 2026-08-27 accessed 2026-09-01
  2. UKGovernmentBEIS/optstop README.md (raw) raw.githubusercontent.com published 2026-08-19 accessed 2026-09-01
  3. Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations (arXiv:2608.14425 abstract) 学術 published 2026-08-14 accessed 2026-09-01
  4. arXiv:2608.14425 full text in HTML (v1) 学術 published 2026-08-14 accessed 2026-09-01
  5. GitHub API: repos/UKGovernmentBEIS/optstop api.github.com accessed 2026-09-01
  6. UKGovernmentBEIS/optstop CHANGELOG.md (raw) raw.githubusercontent.com published 2026-03-18 accessed 2026-09-01