Ai2 TutorMoments: 42 scores across 7 models and 2 prompts
Ai2's TutorMoments preview (7 Aug 2026), from the primary sources: all seven models score 0.223 or below on rigor under a plain prompt; leaders vary by prompt.
This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.
Bottom line: the same model scores very differently under two prompts
On 7 August 2026, Ai2 (the Allen Institute for AI) released a preview of TutorMoments, a framework for measuring whether a model knows when to step in and help a student and when to hold back. According to Ai2, all seven models score 0.223 or below on appropriate rigor under a plain prompt, and every model scores higher on every metric once the prompt spells out the trade-off.
| Item | Detail |
|---|---|
| Published | 7 August 2026 (preview) |
| Where | allenai.org and Hugging Face Blog (same text) |
| Dataset | 462 de-identified US one-on-one maths transcripts, grades 2-7 |
| Annotation | 27 US-based teacher annotators; 122 transcripts annotated |
| Evaluated | 7 models x 3 metrics x 2 prompts |
Every figure and quotation below comes from Ai2’s own publications. As of 10 August 2026 we found no independent third-party verification of TutorMoments.
The full table: 42 scores
These are the values in Table 8 of Ai2’s technical report, matching the tables on the Ai2 blog. The first three numeric columns are the plain prompt, the last three the evaluation-aware prompt.
| Model | Plain: scaf. | Plain: rigor | Plain: no over-scaf. | Aware: scaf. | Aware: rigor | Aware: no over-scaf. |
|---|---|---|---|---|---|---|
| Claude Opus 4.8 | 0.615 | 0.208 | 0.462 | 0.858 | 0.831 | 0.896 |
| Claude Sonnet 4.6 | 0.573 | 0.085 | 0.437 | 0.831 | 0.792 | 0.913 |
| DeepSeek V4 Pro | 0.485 | 0.088 | 0.379 | 0.712 | 0.492 | 0.623 |
| Gemini 2.5 Pro | 0.727 | 0.223 | 0.546 | 0.896 | 0.712 | 0.854 |
| Gemini 3.5 Flash | 0.723 | 0.158 | 0.598 | 0.885 | 0.646 | 0.817 |
| GPT 5.5 | 0.269 | 0.046 | 0.173 | 0.773 | 0.531 | 0.665 |
| GPT 5.4 mini | 0.181 | 0.023 | 0.188 | 0.738 | 0.404 | 0.598 |
The metrics: did the model scaffold when the student needed support, push for rigor when the student was ready for more, and avoid reducing the challenge more than the moment called for.
What the plain prompt actually says
The plain prompt gives only this guidance: “Use everything you know about what makes great tutoring to respond appropriately to the student.” Ai2 writes that under it model tutors often fail to push for rigor and frequently over-scaffold, settling into a “helpful” assistant baseline, and that even GPT 5.5 pushes for rigor appropriately in only 4.6% of cases.
How far the evaluation-aware prompt moves the numbers
Two examples we picked out of Ai2’s table: GPT 5.4 mini goes from 0.023 to 0.404 on rigor, GPT 5.5 from 0.269 to 0.773 on scaffolding. As a ratio the first is roughly 17.6x, but that ratio appears nowhere in the primary sources; we computed it from Ai2’s published values. The source reports absolute numbers.
By condition: the per-metric leader changes
The reordering below is not a sentence Ai2 wrote: it is our re-sorting of the table Ai2 published. Ai2 publishes a separate table per metric and defines no overall ranking.
| Metric | Leader, plain | Leader, aware |
|---|---|---|
| Scaffolding | Gemini 2.5 Pro 0.727 | Gemini 2.5 Pro 0.896 |
| Rigor | Gemini 2.5 Pro 0.223 | Claude Opus 4.8 0.831 |
| Over-scaf. avoided | Gemini 3.5 Flash 0.598 | Claude Sonnet 4.6 0.913 |
The position changes we could confirm span three metrics, five entries.
| Metric | Model | Plain rank | Aware rank |
|---|---|---|---|
| Scaffolding | DeepSeek V4 Pro | 5th | 7th |
| Rigor | Claude Sonnet 4.6 | 5th | 2nd |
| Rigor | Gemini 2.5 Pro | 1st | 3rd |
| Over-scaf. avoided | Gemini 3.5 Flash | 1st | 4th |
| Over-scaf. avoided | Claude Sonnet 4.6 | 4th | 1st |
Ai2 notes that spelling out the trade-off only goes so far: models still differ widely in how they read that prompt.
How to read these scores
- Record which prompt condition a number came from. For a single model, appropriate rigor moves a long way with the condition (GPT 5.4 mini 0.023 and 0.404, Claude Sonnet 4.6 0.085 and 0.792)
- Do not collapse the metrics into one, or conclude from the rigor column alone (see below)
- Read scores next to latency and scope. The strongest models average over 10 seconds per response, on US grades 2-7 maths
How the evaluation works (replay)
According to Ai2, a transcript is paused at a teacher-annotated key moment and a model takes over as tutor. Each replay is five turns in total: three model turns and two synthetic student turns, not five each. The synthetic student is played by Claude Opus 4.6 and is described as an “oracle student” with information the model tutor does not have. An LLM-based pipeline then scores each replay. Ai2 evaluates 520 key moments, split 50/50 between scaffolding- and rigor-appropriate moments.
Caveats
How to handle the human tutor numbers
Ai2 also reports scores for the human tutors in the transcripts: 0.458, 0.182 and 0.496 on the three metrics, obtained with the same pipeline, at the same decision points, five turns after the cut.
In the same passage Ai2 states: “But this isn’t a claim that AI tutors outperform human teachers.” Its reason: annotators specifically looked for moments where tutoring could have gone better, so the dataset concentrates on missed opportunities rather than ideal practice. Ai2 describes human tutors as a naturalistic reference, not a ceiling.
The sources also show an asymmetry: human tutors are scored on the real continuation of the session, models on a replay with a synthetic oracle student. Ai2 notes from its pilot runs that even oracle synthetic students tend to produce unnaturally successful learning behaviours. Same procedure, different counterpart.
Ai2’s own caveat on rigor
Ai2 writes that rigor is noisier than scaffolding, that the pipeline detects rigor pushes less reliably, and that the annotations contain fewer rigor moments (260) than scaffolding ones (738).
Not all 462 transcripts are annotated
The blog mentions 462 transcripts and more than 1,500 teacher-annotated key moments in one sentence, but the technical report states that 122 transcripts are annotated, with 1,536 unique key moments.
The strongest models are the slowest
Ai2 writes that Gemini 2.5 Pro and Claude Opus 4.8, the highest scorers under the evaluation-aware prompt, average over 10 seconds per tutor response, “which may be unrealistic for production deployments in education technology”. That is a possibility, not a verdict.
Limitations Ai2 states
TutorMoments is a preview. Ai2 writes that automated evaluation gives signal about behaviour at a decision point but cannot stand in for studies with real students and learning outcomes; that the dataset is narrow (US-based, mostly elementary and middle-school maths, a single pool of educators); and that findings may not generalise to other subjects or settings.
On why the same score can mean different things, see also the NIST AITE sequestered evaluation programme and how token budgets move agent scores. Figures and quotations were checked against the primary sources on 10 August 2026; see Ai2’s pages for updates.
Sources
- TutorMoments: Do AI tutors know when to help and when to hold back? (Ai2 blog)
- TutorMoments: Do AI tutors know when to help and when to hold back? (Hugging Face Blog)
- When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle (technical report PDF)
- TutorMoments-Preview project site
- allenai/tutormoments-preview (dataset card)
- allenai/tutormoments README (replay pipeline code)
この記事の日本語版: Ai2 TutorMoments: 42 scores across 7 models and 2 prompts(日本語)