ICML 2026 reproductions: 2,226 papers attempted in 19 days
Hugging Face and alphaXiv published the results of the ICML 2026 Open Reproductions challenge on 13 August 2026: 19 days, 1,221 participants, 2,226 papers.
This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.
The headline: 2,226 papers in 19 days, 35,908 claims judged
On 13 August 2026, Hugging Face published a blog post titled “What We Learned by Reproducing 2,200 papers from ICML” (author: abidlabs / Abubakar Abid). It reports the results of the ICML 2026 Open Reproductions challenge, run by Hugging Face and alphaXiv.
According to Hugging Face and alphaXiv, participants brought their own coding agents and reproduced the claims of accepted ICML 2026 papers one claim at a time. The challenge ran from 15 July to 2 August 2026, a period of 19 days.
| Item | Detail |
|---|---|
| Primary source published | 13 August 2026 |
| Organisers | Hugging Face and alphaXiv |
| Period | 15 July to 2 August 2026 (19 days) |
| Scope | Claims extracted from accepted ICML 2026 papers |
| Judging | An automated Logbook Judge, one verdict per claim |
| Verdicts | Frozen in a public dataset at challenge close |
Every figure below comes from the organisers’ own reporting. As of 14 August 2026 we found no independent third-party verification of these totals.
The published numbers (as reported at publication)
| Item | Figure | Unit |
|---|---|---|
| Community members who joined the organisation | 1,221 | people |
| Reproduction logbooks published | 6,816 | logbooks |
| Papers attempted | 2,226 | papers |
| Claims judged | 35,908 | claims |
| HF Jobs launched | 2,962 | jobs |
| Full agent-trace datasets published | 274 | datasets |
| Compute credits per participant | 20 | US dollars |
The 20 US dollars is compute credit per participant for running experiments on HF Jobs, not a monthly fee. The 2,962 counts cloud jobs launched, not participants and not papers. We use 1,221 because that is the figure stated in the blog post itself at the time it was published.
How judging worked: four labels, one verdict per claim
Hugging Face states that judging was performed by an automated Logbook Judge running the open-weights model GLM-5.2, which re-read every logbook and issued a per-claim verdict of one of the following.
- verified
- falsified
- toy (evidence at reduced scale)
- inconclusive
There are four labels. The word contested in the primary source describes a situation, namely independent reproduction teams reaching opposite verdicts on the same claims. It is not listed as a judge label.
That GLM-5.2 is an open-weights model (MIT licence, weights distributed) was confirmed on its model card on 14 August 2026.
Results at a glance: 51% and 23%
Two percentages are published by the organisers.
| Group | Share | Papers |
|---|---|---|
| At least one claim independently verified | 51% | 1,103 |
| At least one claim falsified or contested | 23% | 496 |
In total, 3,978 individual claims were confirmed with real experiments. The primary source summarises the picture as “Reproducibility is not binary; it is adversarial.”
Watch the denominator behind 51% and 23%
Both percentages are stated against “examined papers”. The primary source does not give the number of examined papers, and it is not the same as the 2,226 papers attempted. Writing “51% of the 2,226 papers” would contradict the source, so quote the percentages with the original denominator wording.
How to quote these figures
- Check the unit. Only two figures are per claim: 35,908 judged and 3,978 confirmed. The 1,103 and 496 are counts of papers
- Name the subject. These are self-reported totals from Hugging Face and alphaXiv; no independent verification has been confirmed
- Date the claim. The figures are those stated in the primary source published on 13 August 2026
- Use four labels. Do not present contested as a fifth verdict label
- Go to the frozen dataset. According to the source, verdicts were frozen in a public dataset at challenge close
Caveats
The organisers also record where agents failed
Hugging Face reports that agents got stuck in local loops, misread scale-dependent behaviour, and occasionally built an entire falsification on top of a units mismatch. It adds that the challenge’s most reliable results came from workflows where a human was steering.
Claimed falsifications were re-checked by hand
According to Hugging Face, 35 participants formally claimed they had falsified something, and the organisers adversarially re-verified every claimed falsification. The scope of that re-verification is the falsifications participants formally claimed, not every claim the judge labelled falsified.
On why evaluation conditions change what a score means, see also the NIST AITE sequestered evaluation programme and how compute budgets move agent evaluations. Figures and quotations were checked against the primary sources on 14 August 2026; see Hugging Face’s pages for updates.
Sources
この記事の日本語版: ICML 2026 reproductions: 2,226 papers attempted in 19 days(日本語)