THE PEER-REVIEW LOTTERY
Science runs on peer review: before a result is published, other experts judge it. At the big machine-learning conferences, a handful of reviewers grade each submission and the paper is accepted or rejected. Behind this ritual sits an uncomfortable question. Send the same paper to other reviewers, and would the verdict be the same?
Two expensive experiments
The only direct answers come from two experiments in which a conference reviewed some papers twice, with two independent committees. The paper recalls their results:
- NeurIPS 2014: of 166 doubly reviewed papers, the two committees disagreed on 43 — 25.9 percent. About half of the papers accepted by one committee would have been rejected by the other.
- NeurIPS 2021: on 882 papers, the committees disagreed on 23.0 percent.
Such experiments are rare because they are costly: they mean doubling the reviewing of hundreds of papers. Feilian Huang, an independent researcher, asked whether the same number could be estimated from public data alone.
Simulating a second committee
The study uses the public reviews of the ICLR conference from 2017 to 2025: 36,113 papers and 134,912 reviews. A statistical model splits each score into two parts: the hidden quality of the paper, and the “noise” added by each reviewer. A second model reproduces how scores turn into decisions. The computer then imagines, for each paper, two independent committees of two, three or four reviewers, a thousand times over, and counts how often they disagree.
Before trusting the result, the author checked it twice. Against the NeurIPS 2021 experiment, with three reviewers per paper as there, the simulation gives 23.3 percent disagreement for 23.0 percent measured. And on ICLR papers with at least four reviews, simply splitting the real reviewers into two random pairs gives the same answer as the model, within one point, in 2018 and in every year from 2021 to 2025.

Share of score variance due to the paper itself, year by year: around half, the rest is reviewer noise. — Figure 2, Huang (2026), arXiv:2610.06591.
About one paper in four
The share of score variation that comes from the paper itself ranges from 0.38 to 0.58 depending on the year — the rest is reviewer noise, in line with the 2014 estimate of about half. The consequences:
- With two reviewers per committee, the decision flips for 22.8 percent (2017) to 30.3 percent (2025) of papers. With four reviewers, 17.7 to 24.2 percent.
- 30 to 50 percent of accepted papers would have been rejected by an independent committee — up to half in 2021.
- Over nine years in which submissions grew about 24-fold (489 in 2017, 11,663 in 2025), the study finds no robust trend: the noise neither clearly grows nor shrinks. The author stresses that with only nine years of data, this is evidence against a strong trend, not proof of stability.

Simulated disagreement between two committees, compared with the two real NeurIPS experiments. — Figure 3, Huang (2026), arXiv:2610.06591.
The scale itself adds chance
Two years stand out. In 2020, ICLR used a coarse four-level grading scale (1, 3, 6, 8): 23.7 percent of papers received identical scores from all their reviewers, against 4.7 to 13.1 percent in other years. To test the effect, the author took the 2021 scores, given out of 10, and recoded them on the four-level scale. Result: about 8 percent of the signal is lost, disagreement between committees rises by about 7 points, and the share of accepted papers that would flip rises by about 11 points. A coarse scale does not just relabel the same opinions; it adds chance.
2021 is different: a fine scale, but the weakest signal of the nine years, and 30.4 percent of accepted papers sat just above the acceptance threshold. Many papers near the line, plenty of noise: easy flips.
And chatbots?
The study planned to look for a change in reviewing behaviour after 2023, when language models spread. The public data contain no review text or reviewer confidence after 2022, and the statistical test available has almost no power. The author therefore draws no conclusion either way — and presents that restraint as deliberate.
Auditing the lottery
The method needs only scores and decisions, which several conferences already publish. Any of them could compute its own “lottery” rate without duplicating a single review. The author’s verdict on nine years of ICLR: reviewing is neither collapsing nor improving — it stays noisy at about the same level. The study covers ICLR only, and its author is a single independent researcher.
