optstop 57-97% savings: the denominator behind the number
optstop reports 57-97% savings on planned trials. The range comes from nine shadow-mode settings, and the README states where its intervals are calibrated.
This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.
Conclusion: the 57-97% figure is a counterfactual, not a live result
On 27 August 2026 the UK AI Security Institute (formerly the AI Safety Institute) published a blog post introducing optstop, an open source tool that stops LLM evaluation sampling on statistical criteria. The headline range is 57% to 97% of planned trials. Three things define what that range is.
- According to the accompanying paper (arXiv:2608.14425, by Toby D. Pilditch), the range was measured in shadow mode: every planned trial is still executed, while the run internally records the point at which the stopping criteria would have been met
- The denominator is “planned trials across nine validation settings”, tied to an illustrative 200-item, 10-epoch evaluation design
- The README, not the blog, states where the credible intervals are well calibrated: 50 or more items per grouping, with performance in the 0.2-0.8 range
Every number below comes from the UK AISI and the paper’s author. As of 1 September 2026 no independent third-party replication or verification has been found.
The three primary documents word the savings differently
| Document | Date | How the savings are stated | Caveat carried in the same document |
|---|---|---|---|
| AISI blog | 2026-08-27 | Two sentences: “under every condition we set … 57% and 97% of planned runs” and “under every condition … 57% and 97% of planned trials” | With a leaner design and fewer trials, the number of trials saved may be lower |
| Paper abstract | 2026-08-14 | “removes 57%-97% of planned trials across nine validation settings” | “with the magnitude of savings depending on evaluation design” |
| README | 2026-08-19 | No savings figure; calibration conditions instead | For small samples or near-boundary performance, “CIs may undercover” |
The blog’s two sentences are not identical. The first carries the qualifier “we set” and refers to planned runs; the second drops the qualifier and refers to planned trials.
Measurement conditions behind the range
The paper reports a 3x3 matrix experiment crossing three inference pathways (binary, ordinal, continuous) with three performance levels.
| Item | Value |
|---|---|
| Validation settings (cells) | 9 |
| Items per cell | 200 |
| Epochs per item | 10 |
| Planned trials per cell | 2,000 |
| Precision threshold δ | 0.05 |
| Credible interval level | 97% |
| Mean efficiency gain | 81.1% |
| Mean absolute score deviation | 0.006 |
One cell differs: mid-binary uses the full GPQA Diamond set, so it has 198 items and 1,980 planned trials. The 0.006 deviation is a mean absolute value on a normalised [0,1] scale, comparing full-run and truncated estimates. All nine cells triggered early stopping.
Stopping rules, and what happens when neither fires
Per the blog, sampling stops when either rule is satisfied.
- Precision: the credible interval is narrow enough
- Stabilisation: the interval has stopped changing, so the data has nothing more to tell us
If neither rule is met, the grouping simply runs to its full planned budget. A genuinely noisy estimate is not cut short; it is sampled to the end. A conservatism mechanism also demands more data when estimated success falls below 1% — that threshold is the default in version 0.4.0, whose CHANGELOG entry is dated 18 March 2026.
Adoption path described by the blog
- Post-hoc: run it over a completed evaluation and confirm the stopped estimates match the full-run estimates
- Shadow mode: run the evaluation as normal while optstop reports the saving it would have made, with the full dataset preserved
- Live: switch on live stopping once the shadow-mode numbers are convincing
The package ships an OptimalStoppingManager implementing the EarlyStopping protocol of inspect_ai, the UK AISI evaluation framework; the shadow_mode parameter defaults to False. The blog adds one practical requirement: tasks must be presented in a randomised order so that early estimates are not biased by queue ordering, and it states that Inspect supports this.
Caveats
Intervals are called well-calibrated only within a stated range
| Condition | What the README says |
|---|---|
| 50+ items per grouping, performance 0.2-0.8 | Credible intervals are well calibrated |
| Smaller samples, or performance close to 0% or 100% | Intervals may undercover, due to hierarchical shrinkage |
| Stress test: 18-30 items per grouping with many boundary performers | Observed coverage can fall substantially below nominal |
Only the third row is stated as observed; the second is a possibility. The README frames this as a structural property of Bayesian hierarchical models rather than an implementation bug, and notes it affects live and post-hoc modes equally.
27 August 2026 is the date it was written about, not released
| Event | Date |
|---|---|
| Repository created (GitHub API metadata) | 2025-12-12 |
| Latest CHANGELOG entry, 0.4.0 | 2026-03-18 |
| README commit “Initial public release” | 2026-06-12 |
| README last updated | 2026-08-19 |
| Introduced on the AISI blog | 2026-08-27 |
The current version is 0.4.0 and the licence is MIT, both as retrieved on 1 September 2026.
Related articles
- UK AISI: How Token Budgets Change AI Agent Benchmark Scores — the same institute on how the compute you grant an agent changes what a score means
- NIST AITE: Sequestered AI Evaluation and Its Three Tasks — why identical scores are not comparable across measurement setups
Sources
- Optimal stopping: spending evaluation compute where it counts (AISI Blog)
- UKGovernmentBEIS/optstop README.md (raw)
- Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations (arXiv:2608.14425 abstract)
- arXiv:2608.14425 full text in HTML (v1)
- GitHub API: repos/UKGovernmentBEIS/optstop
- UKGovernmentBEIS/optstop CHANGELOG.md (raw)
この記事の日本語版: optstop 57-97% savings: the denominator behind the number(日本語)