Last verified 4 min read AI evaluation and benchmarks

ICML 2026 reproductions: 2,226 papers attempted in 19 days

Hugging Face and alphaXiv published the results of the ICML 2026 Open Reproductions challenge on 13 August 2026: 19 days, 1,221 participants, 2,226 papers.

This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.

The headline: 2,226 papers in 19 days, 35,908 claims judged

On 13 August 2026, Hugging Face published a blog post titled “What We Learned by Reproducing 2,200 papers from ICML” (author: abidlabs / Abubakar Abid). It reports the results of the ICML 2026 Open Reproductions challenge, run by Hugging Face and alphaXiv.

According to Hugging Face and alphaXiv, participants brought their own coding agents and reproduced the claims of accepted ICML 2026 papers one claim at a time. The challenge ran from 15 July to 2 August 2026, a period of 19 days.

ItemDetail
Primary source published13 August 2026
OrganisersHugging Face and alphaXiv
Period15 July to 2 August 2026 (19 days)
ScopeClaims extracted from accepted ICML 2026 papers
JudgingAn automated Logbook Judge, one verdict per claim
VerdictsFrozen in a public dataset at challenge close

Every figure below comes from the organisers’ own reporting. As of 14 August 2026 we found no independent third-party verification of these totals.

The published numbers (as reported at publication)

ItemFigureUnit
Community members who joined the organisation1,221people
Reproduction logbooks published6,816logbooks
Papers attempted2,226papers
Claims judged35,908claims
HF Jobs launched2,962jobs
Full agent-trace datasets published274datasets
Compute credits per participant20US dollars

The 20 US dollars is compute credit per participant for running experiments on HF Jobs, not a monthly fee. The 2,962 counts cloud jobs launched, not participants and not papers. We use 1,221 because that is the figure stated in the blog post itself at the time it was published.

How judging worked: four labels, one verdict per claim

Hugging Face states that judging was performed by an automated Logbook Judge running the open-weights model GLM-5.2, which re-read every logbook and issued a per-claim verdict of one of the following.

  • verified
  • falsified
  • toy (evidence at reduced scale)
  • inconclusive

There are four labels. The word contested in the primary source describes a situation, namely independent reproduction teams reaching opposite verdicts on the same claims. It is not listed as a judge label.

That GLM-5.2 is an open-weights model (MIT licence, weights distributed) was confirmed on its model card on 14 August 2026.

Results at a glance: 51% and 23%

Two percentages are published by the organisers.

GroupSharePapers
At least one claim independently verified51%1,103
At least one claim falsified or contested23%496

In total, 3,978 individual claims were confirmed with real experiments. The primary source summarises the picture as “Reproducibility is not binary; it is adversarial.”

Watch the denominator behind 51% and 23%

Both percentages are stated against “examined papers”. The primary source does not give the number of examined papers, and it is not the same as the 2,226 papers attempted. Writing “51% of the 2,226 papers” would contradict the source, so quote the percentages with the original denominator wording.

How to quote these figures

  1. Check the unit. Only two figures are per claim: 35,908 judged and 3,978 confirmed. The 1,103 and 496 are counts of papers
  2. Name the subject. These are self-reported totals from Hugging Face and alphaXiv; no independent verification has been confirmed
  3. Date the claim. The figures are those stated in the primary source published on 13 August 2026
  4. Use four labels. Do not present contested as a fifth verdict label
  5. Go to the frozen dataset. According to the source, verdicts were frozen in a public dataset at challenge close

Caveats

The organisers also record where agents failed

Hugging Face reports that agents got stuck in local loops, misread scale-dependent behaviour, and occasionally built an entire falsification on top of a units mismatch. It adds that the challenge’s most reliable results came from workflows where a human was steering.

Claimed falsifications were re-checked by hand

According to Hugging Face, 35 participants formally claimed they had falsified something, and the organisers adversarially re-verified every claimed falsification. The scope of that re-verification is the falsifications participants formally claimed, not every claim the judge labelled falsified.

On why evaluation conditions change what a score means, see also the NIST AITE sequestered evaluation programme and how compute budgets move agent evaluations. Figures and quotations were checked against the primary sources on 14 August 2026; see Hugging Face’s pages for updates.

Sources

  1. What We Learned by Reproducing 2,200 papers from ICML (Hugging Face blog) huggingface.co published 2026-08-13 accessed 2026-08-14
  2. ICML 2026 Agent Reproductions (organising account profile) huggingface.co accessed 2026-08-14
  3. ICML-2026-agent-repro/challenge (dataset of frozen verdicts) huggingface.co accessed 2026-08-14
  4. zai-org/GLM-5.2 model card (MIT licence, weights distributed) huggingface.co accessed 2026-08-14