NIST AITE: Sequestered AI Evaluation and Its Three Tasks
NIST's AITE evaluates VLMs on data that is never released. Trial counts (641, 10,000, 3,000), the real metric definitions, and how results are published.
This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.
The short answer: three tasks, non-public data, evaluation period starting August 2026
The Technology Test and Evaluation Division at NIST announced the Artificial Intelligence Technology Evaluation (AITE) in a news item dated July 27, 2026 (the page reads “Released July 27, 2026, Updated July 28, 2026”). The three initial tasks all involve image analysis with large vision language models (VLMs).
The difference from a public benchmark is that the test data is never released. NIST describes AITE as volunteer testing of AI models on blind data, where the sequestered environment “mitigates the risk of train/test data contamination.” Note that the program’s list of objectives uses a stronger verb — to “remove the potential for train/test data contamination.” The two phrasings are not identical, and the softer one should govern how you read the numbers.
The schedule is published only at month granularity. As of August 3, 2026:
| When | What |
|---|---|
| July 2026 | AITE Kickoff |
| July 2026 | Evaluation Plan Release |
| August 2026 | Evaluation Period Begins |
No specific dates or deadlines appear on the public pages. The Evaluation Plan itself is the “NIST AI Technology Evaluation Overview & Road Map” v1.0, dated July 24, 2026.
Task specifications, side by side
Everything below comes from NIST’s own published specifications — the Available Tests table in the AITE Overview and each task’s Test Specification v1.0. As of August 3, 2026 we found no third-party reporting of these figures.
| Task (dataset) | Trials | Data size |
|---|---|---|
| Quantum Dot Control (QDC Patches v1.0) | 641 | ~0.5 GB |
| Human Genome Variant Curation (Genome Variant Visualization v1.0) | 10,000 | ~1.1 GB |
| Public Safety Visual Event Recognition (Gumby V1.0) | 3,000 | ~9 GB |
The 9 GB for Public Safety covers 9,000 images grouped in sets of three per potential event, which is how 3,000 trials arises. The Test Specifications are dated May 28-29, 2026, roughly two months before the announcement, and all three datasets are listed as created in April 2026.
Three metric names, but two of them are the same function
Read only the Overview table and the three tasks appear to use unrelated metrics. Open the Test Specifications and two of them turn out to be the same detection cost function with different settings.
| Task | Metric name in the Overview table | Definition in the Test Specification |
|---|---|---|
| Quantum Dot Control | Mean Squared Error | The mean of the mean squared error computed across dimensions |
| Human Genome Variant Curation | Average Error Rate | An unweighted detection cost function |
| Public Safety Visual Event Recognition | Detection Cost Function | A weighted detection cost function |
The genomics specification defines its primary metric as “an unweighted detection cost function, consisting of the average of the Type I and Type II error rates,” i.e. C = (P_Miss + P_FalseAlarm) / 2, while the public safety specification states that “a weighted detection cost function will be computed.” So these two metrics are not different families. Quantum Dot Control is not plain MSE either: its primary metric is “the mean of the mean squared error computed across dimensions.”
Input formats differ too
| Task | Input | Output |
|---|---|---|
| Quantum Dot Control | Single image and text | Text (three-state probability estimation) |
| Human Genome Variant Curation | Image of a genome region plus text describing variants | Whether all given variants were correctly called |
| Public Safety Visual Event Recognition | Three video key frames and text | A yes or no decision |
The Overview table lists modalities as “text and image / text” for all three, but the Public Safety Test Specification gives the input as “Multiple Images and Text” and states that the test data consists of key frames selected from videos. The three tasks do not share one input format.
Human Genome Variant Curation is a VLM evaluation task — deciding whether variants shown in an image were called correctly — not a diagnostic or clinical task. Its reference labels are described as based on highly accurate assembly from multiple sequencing technologies and confirmed by experts.
The two participation tracks
| Track | What you submit | What you receive |
|---|---|---|
| Data provider | An original dataset in your domain that is inaccessible to others, plus a meaningful task on it | Careful measurements of top models performing your task on your data |
| Model provider | AI models to be tested on the datasets and tasks | How your models perform relative to others on the same data using the same metrics |
The steps described on the public pages are limited to the following.
- Confirm you can abide by the AITE Participation Agreement and rules. Participation is described as open to all who wish to engage in one of the ways above.
- Contact aite-poc@list.nist.gov to request participation or ask questions.
- Wait for intake. In the Overview & Road Map v1.0, NIST states that the initial phase will accept “a very limited number” of additional tests and models, that data providers are selected on a combination of anticipated ease and impact, and that submitted models are queued and tested on a first come, first served basis, as capacity allows.
Submitting a request therefore does not guarantee evaluation. The AITE Platform and Participant Login links were still marked “Coming soon” as of August 3, 2026.
How results are published
This is spelled out in the public documents. According to the NIST Overview, results for all submitted systems are posted on the AITE website along with the identification of the submitting organizations, and NIST summary reports are updated “no less than once a year.”
Section 3 of the Overview & Road Map v1.0 adds that NIST will publish and periodically update a leaderboard and other reports, and that the displayed results will specify, for each eligible model, the following five fields.
- dataset
- task
- model by metric
- scores
- measures of uncertainty
The public documents do not define what makes a model “eligible.”
Caveats
What the public pages do not state, as of August 3, 2026
Across the four pages checked on that date — the news item, AITE Main, the Overview, and the Overview & Road Map v1.0 — there is no statement about nationality or location requirements for participants, and none about fees. The text of the Participation Agreement is not linked from the public pages either.
In other words, the pages neither grant nor deny eligibility to any particular country’s organizations. Absence of a requirement in public documentation is not proof that no requirement exists. Anyone considering participation should contact aite-poc@list.nist.gov.
The program is still small
The Overview & Road Map v1.0 describes the current state of AITE as consisting of three tests and results from a single model, developed in four phases, and states that the precise scope and timing of each phase will be contingent on capacity and levels of engagement. The Overview’s model listing likewise contained a single entry as of August 3, 2026.
For other cases of opening a single summary-table phrase down to the primary source, see whether a multilingual retrieval model was trained on Japanese and which tokens the npm 2FA-bypass restriction actually covers. Check the official NIST pages for the current status and for anything the published documents do not cover.
Sources
- Announcing NIST's Artificial Intelligence Technology Evaluation (AITE)
- AITE Overview — AI Technology Evaluation (NIST Pages)
- AITE Main / Schedule (ai-challenges.nist.gov)
- NIST AI Technology Evaluation Overview & Road Map v1.0
- NIST AITE Test Specification: Quantum Dot Control v1.0
- NIST AITE Test Specification: Human Genome Variant Curation v1.0
- NIST AITE Test Specification: Public Safety Visual Event Recognition v1.0
- NIST unveils new AI evaluation platform (Nextgov/FCW)
- NIST Launches New Testbed Program to Evaluate AI Model Performance (ExecutiveGov)
この記事の日本語版: NIST AITE: Sequestered AI Evaluation and Its Three Tasks(日本語)