NAAIMES best practice: automated evaluation only, 6 excluded
The first NAAIMES best practice document (23 July 2026) covers only automated evaluations of LLMs; six evaluation types are explicitly out of scope.
This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.
Conclusion: the scope is automated evaluation of large language models
The first best practice document from NAAIMES (International Network for Advanced AI Measurement, Evaluation and Science, formerly the International Network of AI Safety Institutes) is titled “Best Practice: Automated Evaluation of Large Language Models”. The UK AI Security Institute (AISI) announced its completion in a blog post dated 23 July 2026. The wording is “completion”, not the publication of a standard.
The document itself states: “This document focuses on one important subset: automated evaluations of AI models.” The UK AISI blog post never mentions this restriction, so readers who stop at the blog will overestimate the scope.
The Network was established in November 2024 and brings together Australia, Canada, the European Union, France, Japan, Kenya, the Republic of Korea, Singapore, the United Kingdom and the United States - ten countries and regions, since the EU is not a country. NIST listed the same set as of February 2026.
What is in scope and what is not
| Category | Item |
|---|---|
| In scope | Automated evaluations of AI models |
| Out of scope | Red-teaming and safeguard testing |
| Out of scope | Agentic evaluations |
| Out of scope | Expert capability evaluations |
| Out of scope | Propensity evaluations |
| Out of scope | Human uplift studies |
| Out of scope | Open-world evaluations |
| Not covered | How to build your own evaluations; interpretation of results |
On the six excluded areas, the document says only that they “may be addressed in future publications by the Network”. That is a possibility, not a commitment.
Position: complementing NIST AI 800-2
The document is aimed at third-party evaluators. Both the blog and the document state that it builds on and is intended to complement existing best practice publications, most notably NIST AI 800-2. It neither claims conformance nor replaces them. The reference list describes NIST AI 800-2 as an “Initial Public Draft”, so as of July 2026 it is not a finalised standard.
The four related posts, by date and by author
| Publisher | Date | Subject |
|---|---|---|
| CAISI (Canada) | 2026-07-08 | What information evaluators should share |
| Singapore AISI | 2026-07-08 | How to test AI systems rather than models |
| INESIA (France) | 2026-07-08 | What to prioritise when resources are limited |
| EU AI Office Safety Unit | not stated | Primary source not reached; contents not discussed here |
| NAAIMES / UK AISI | 2026-07-23 | The best practice document and the summary blog |
The three posts came first; the document followed about two weeks later. Calling them simultaneous does not match the dates.
The series has four posts because the CAISI page itself states that it “also includes posts from: The European Union (EU) AI Office Safety Unit”, INESIA and the Singapore AI Safety Institute. The UK AISI blog simply covers three of them, so “three countries published three posts” contradicts the primary source.
These are not Network conclusions
UK AISI presents them under the heading of addressing open questions. Each post is authored by an individual member, and no Network-level answer is given. As of February 2026, UK AISI had listed five open questions; the titles of the three July posts correspond to three of them. None of the three corresponds to the remaining two, but because the EU post exists, those cannot be called unanswered.
Steps: applying the published guidance
- Check that your evaluation is an automated one. Agentic evaluations and red-teaming fall outside this document
- Fix the scoring side. According to the document (p.7), the evaluation of outputs “should always be comparable”
- Set the generation side by purpose, and record the choice. The same document says how strictly generation is standardised “depends on the goals of the evaluation”
- Split your data if you run capability elicitation. The proportions are below
- Decide in advance what to report and what to withhold. The CAISI list of five items and three categories is the material
Data split for capability elicitation (document p.9)
| Purpose | Amount | Unit |
|---|---|---|
| Exploratory set (manual experimentation, initial qualitative analysis) | 2-5 | items |
| Tuning set (iterating over prompts and parameters) | 70-80 | % |
| Evaluation set (final reported evaluation only) | 20-30 | % |
The “2-5” is a count of items; the other two are shares of the items. The three sets should be disjoint and evaluators should not look at the evaluation set during tuning; the document calls failing to observe this “one of the most common and consequential mistakes in capability elicitation”. A footnote adds that the percentages are a rule-of-thumb.
CAISI: five items to report, three categories to withhold
| Category | Item |
|---|---|
| Report | Intent of the measurement and the contexts of use |
| Report | Methodology: what is measured and how |
| Report | The model or system tested, and its settings |
| Report | Caveats and limitations, including uncertainty statistics |
| Report | Potential conflicts of interest in producing and reporting results |
| Withhold | Datasets containing safety critical information (CBRN and cybersecurity are the examples given) |
| Withhold | Adversarial testing insights that could provide uplift to malicious actors |
| Withhold | Test sets that risk data contamination |
Disclosure is not binary: the page distinguishes public dissemination, trusted actors, and developers or system owners.
Caveats
The same page carries two different lists of five
The CAISI page has two bulleted lists of five, and they differ. The fifth item under “Key takeaways” is “Practicing selective disclosure”, whereas the table above follows the list introduced by “At minimum, we propose that evaluators should”. Do not merge them into six. The only date on the page is a “Date modified” value.
Singapore’s process has four steps and is not its own invention
Singapore AISI cites IMDA’s “Starter Kit for Testing LLM-based Applications for Safety and Reliability” and writes “We present an adapted version here”, followed by four steps: identifying relevant risks and calibrating the extent of testing; constructing a realistic testing environment; running evals by constructing test datasets, metrics and evaluation techniques; and analysing results and trajectories to inform improvements and mitigations. “Establishing fit-for-purpose thresholds” is a separate callout, not a fifth step. Trajectories here means execution traces, not tendencies.
INESIA distinguishes may from should
The INESIA framework, published on the SGDSN site, has five elements: decision context, scenarios and threat models, test design and escalation, reporting, and ecosystem. The original calls it “a minimal framework”. Starting with lightweight probes or smoke tests is permissive - “evaluators may begin with lightweight probes” - while “more demanding tests should be prioritised where the stakes are higher” is an obligation form.
Nothing in the sources says the guidance is binding
The document uses the recommendation form “evaluators should…”, and no obligation wording such as must, mandatory, binding or compliance appears. No official statement calls it voluntary either, so this is best treated as an absence of any statement. Translated versions are not mentioned, and as of 21 August 2026 only the English version is available. On future work, the blog says the Network “will align on further best practices […] with particular attention given to agentic capabilities” - a plan, not an agreement already reached.
Japan is a member of the Network but is not among the authors of this series. The abbreviation varies too: the Canadian page writes “AAIMES Network” where the others write NAAIMES.
The numbers and quotations here were checked against the primary sources on 21 August 2026. Please confirm any updates on the official pages.
Sources
- International evaluation best practice and open questions in AI measurement (UK AISI blog)
- International Network for Advanced AI Measurement, Evaluation and Science Best Practice: Automated Evaluation of Large Language Models (PDF)
- What information should AI evaluators share? (Canadian AI Safety Institute, CAISI / ISED)
- How should evaluators test AI systems (as opposed to models)? (Singapore AI Safety Institute)
- What should evaluators prioritise when resources are limited? (INESIA / SGDSN, France)
- International consensus and open questions in AI evaluations (UK AISI, the five open questions of February 2026)
- International Network for Advanced AI Measurement, Evaluation, and Science Publishes Consensus Areas on Practices for Automated Evaluations (NIST news)
この記事の日本語版: NAAIMES best practice: automated evaluation only, 6 excluded(日本語)