What We Learned by Reproducing 2,200 papers from ICML

2026-08-20 · Hugging Face

What We Learned by Reproducing 2,200 Papers from ICML

The Scale Problem in AI Research

Questions about the reproducibility of AI research predate the current wave, but they are dramatically worsened by scale. ICML 2026 received 23,918 submissions and accepted 6,352 papers — roughly double the previous year. This exponential growth is partly driven by AI agents that make it faster to run experiments and write papers.

Reviewing capacity has not kept pace. Reviewers are mostly volunteers who often lack the time or expertise for thorough evaluation. A reviewer of one accepted ICML 2026 spotlight paper wrote: "My low confidence score is because I did not check all the proofs carefully."

The same technology driving the paper flood can also help solve the problem. Coding agents like Claude Code, Codex, Cursor, and Pi can read a paper, implement the code, run experiments, and report findings. What once cost a reviewer an entire weekend can now be attempted by an agent in an afternoon — and thousands of agents can work in parallel.

The central question was: if we re-examined a major conference at scale and tried to reproduce every paper, what would we discover?

The Hackathon (July 15 – August 2, 2026)

Instead of auditing papers internally, the challenge was opened to the entire community, bringing diverse agent frameworks, compute budgets, and scientific tastes.

How It Worked

  • Pick a paper: All 6,341 accepted papers were indexed with abstracts and their core scientific claims were automatically extracted. Agents started from concrete, checkable targets rather than 40-page PDFs. Multiple reproductions of the same paper were encouraged.
  • Bring your own agent: Participants used Claude Code, Codex, Cursor, OpenResearch's orx, and many others. A streamlined interface allowed an agent to fetch the paper, claims, and instructions with a single command.
  • Reproduce and publish everything: Every run produced a Trackio logbook — a static Hugging Face Space containing the write-up, executed code, artifacts, and optionally the full agent execution trace uploaded as a Hugging Face Dataset. The auditing process itself had to be fully auditable.
  • Get judged: An automated Logbook Judge powered by the open-weights GLM-5.2 model re-read every logbook and issued per-claim verdicts (verified, falsified, toy, or inconclusive). The judge was instructed to treat each logbook's self-assessment as untrusted.

Participants received $20 in Hugging Face compute credits. Across the challenge, 2,962 cloud jobs were launched. When full reproduction was impossible (proprietary datasets or unreleased checkpoints), participants ran toy reproductions on synthetic data that mimicked the original properties.

Key Numbers

  • 1,221 community members participated
  • 6,816 reproduction logbooks published
  • 2,226 papers attempted (34% of the entire conference), many by multiple independent teams
  • 35,908 claims judged, with all verdicts frozen in a public dataset
  • 2,962 HF Jobs launched and 274 full agent-trace datasets published

What We Found

Aggregating claim-level verdicts per paper:

  • 51% of examined papers (1,103) had at least one claim independently verified. Of these, 266 papers were fully reproduced (every claim verified), and 632 more were partially reproduced with nothing falsified. In total, 3,978 individual claims were confirmed with real experiments.
  • 23% of examined papers (496) had at least one claim falsified or contested. This includes 49 papers where all claims were falsified, and notably 242 papers where independent teams reached opposite verdicts on the same claims. Reproducibility is not binary — it is adversarial.

The remainder were in the middle: 502 papers with only toy-scale evidence and 280 where nothing could be established (most commonly due to missing artifacts).

Reproductions Done Well

Some papers emerged looking excellent. Notable examples include:

  • "Flat Minima and Generalization: Insights from Stochastic Convex Optimization" was reproduced by 20 independent teams, 12 of which verified every claim. One reproduction published the full agent trace.
  • "A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness" had 14 of 17 logbooks verify every claim. A paper about unreliable LLM judges that itself held up under scrutiny by LLM agents.

Falsifications

35 participants formally claimed they had falsified something. The team began adversarially re-verifying these claims.

Source