What this is
The project
An attempted description of peer review, at the resolution of a single thought
The Atlas of Judgment is a metascience instrument: an analysis of every public peer review of ICLR — the International Conference on Learning Representations — asking what the act of reviewing is. To make that analysable, each review is read into atomic units of evaluative logic — and what kinds of unit the record turns out to contain is itself part of the question — then the units are counted, compared, and tested at scale. Not what did reviewers decide, but how did they think.
Peer review decides what counts as knowledge in machine learning. It runs on the unpaid attention of tens of thousands of people, it is argued about constantly, and its interior is studied far less than its verdicts. What is visible from outside is the output — a score, an accept, a reject — the way a verdict is visible while the deliberation behind it is not. This project is an attempt to open the deliberation and describe it.
The result is 1,420,178 atomic units of reviewer reasoning, drawn from nine years of ICLR and analysed across thirty-three plates. It is not an argument about peer review. It is an attempt at a description of it — made in one chosen form, carried all the way through — and it was built so that any part of the description can be taken apart by someone who doubts it.
792,703 messages
The whole project in one strip: the public record, the four-part unit it is dissected into, the 12 × 12 grammar each unit is filed under, and the plates that count them. Read backwards, it is the route of doubt: any number on any plate descends, step by step, to the reviewer sentences beneath it.
The motive
A score is a judgment with the reasons thrown away
Two reviewers write 5. One means the idea is not new. The other means the experiments do not support the claim. Averaged into a panel statistic, the difference disappears — and the difference was the whole of what happened. A score is a compression, and what it compresses away is precisely the part that could be argued with, learned from, or improved.
The project's research question was fixed early, and verbatim, in its design documents: not whether a paper was accepted, nor whether a review was favorable, but — what do reviewers inspect, and through which criteria, comparisons, assumptions, and inferences does that observation become a judgment?
Peer review is a central quality-control mechanism of science, yet the reasoning inside it is rarely studied as an object in its own right. Reviews are usually reduced to scores, sentiment, or accept/reject outcomes. This project treats the reasoning itself as the data.
That choice sets the shape of everything downstream. Sentiment analysis would have counted how reviewers felt. Outcome modelling would have predicted what became of the paper. Neither says what was being weighed. The atlas is built for a narrower and more stubborn question: which part of a paper was looked at, and which standard was held against it. Asking that at scale meant choosing a grain and a vocabulary before any answer could exist — the four-part unit, the twelve-by-twelve grammar — so every answer the atlas gives is an answer in that chosen vocabulary: one way of carving the reasoning, carried through, not the only way it could have been carved.
The venue
The venue that publishes the reviews of the papers it rejects
One property makes ICLR nearly unique among top machine-learning venues: it publishes the reviews of rejected papers. Public review archives usually begin where acceptance ends — and a record without the rejected majority shows judgment only where judgment said yes. ICLR's OpenReview forums carry the full distribution — every submission, every review, every rebuttal, every verdict — for nine consecutive years, 2018 through 2026. That is 52,460 submissions and 792,703 forum messages, all from a single source with a single collection path.
One forum’s record, drawn to its parts. The counts are the corpus’s own: ≈3.8 official reviews per paper and ≈425 words per review are the 2026 means (74,380 reviews over 19,474 papers); 52,460 submissions and 792,703 messages are the nine-year totals. The fork’s two ends are drawn with the same ink on purpose — the record’s completeness does not depend on the verdict, and that is the property this atlas rests on.
The choice is a constraint, not an endorsement. ICLR is one venue, in one field, with one culture and a review form that changed four times across the nine years. Everything in the atlas describes ICLR; nothing in it should be read as a fact about peer review in general unless the plate says so, and none of them do.
The atom
One complete movement of a reviewer's mind
The unit of analysis is what this project treats as one complete movement of a reviewer's mind — from a thing inspected, through an observation, through the standard invoked, to a judgment. A real unit from the corpus:
Specimen № 335 of 410,586, from ICLR 2026 submission hPUjTTj64O. The same unit is opened part-by-part on the atlas's own Plate 0.
Each unit also carries a verdict polarity (negative / positive / conditional / uncertain / mixed), a suggested fix when the reviewer offered one, line-level citations back to the source review, and a flag recording whether it is grounded in the reviewer's explicit words or inferred by the analysis layer.
Everything in the atlas is a count of these. No plate holds anything that is not, at bottom, a pile of units — which is why every figure can be opened until a sentence written by an actual reviewer is showing. A number here is a door, not a conclusion.
The making
Three layers, kept separate on purpose
No one reads 792,703 forum messages. The scale that makes the question answerable is exactly the scale that puts a machine between the reader and the reviewer. The honest way to handle that is not to hide the machine but to name the layers and keep them apart, so that a claim can always be traced to the layer that produced it.
Two tracks, one taxonomy
Two independent extraction tracks cover complementary views. The 2026 track works review-by-review at high resolution: 74,380 reviews → 410,586 units. The 2018–2026 track works forum-by-forum across nine years: 50,861 forums → 1,009,592 units, with temporal fields (before/after rebuttal, judgment change) that power the drift and rebuttal analyses.
Both tracks share one taxonomy — 12 objects of scrutiny × 12 reasoning standards — induced from the data itself: a 12,000-unit sample was embedded and clustered with no seeding from prior schemes, the clusters were read and named by hand, and every unit in both tracks was then assigned to its nearest category. Definitions for every term are in the atlas lexicon, and on hover throughout.
The limits
A machine reading of reviews, not a measurement of reviewers
What this is not: a direct measurement of human reviewers. Every unit is a two-stage machine reading. The support_status flag (explicit vs. inferred) and a per-category reliability grid in the atlas appendix exist so that every claim can be read against the instrument's own error.
Read the atlas with five caveats. (1) It measures a machine reading of reviews, not reviewers directly; the reliability grid shows where inference substitutes for quotation. (2) The negative share reflects the memo prompts' bias toward articulating criticism — it is not a sentiment ratio. (3) Coverage is 98.1–98.2% on both tracks; the missing tail (parse failures, provider content-policy refusals) is documented and its topic correlation unverified. (4) Decision-linked plates (the Consequence) are associations: outcomes were never shown to the extraction pipeline, but criticism breadth and paper quality are entangled by design. (5) The four-part unit and the 12 × 12 taxonomy are one induced carving of the text, not the only possible one: their internal reliability is measured, and the categories track ICLR's own sub-scores where they should, but no check can establish them as the right carving — a different scheme would draw a different map of the same reviews.
A caveat kept in an appendix is a caveat that has been hidden. These ride at the same weight as the findings, on the plates themselves as well as here — and the atlas's third appendix is given over entirely to results that failed: patterns that looked real, were tested against a null, and did not survive it.
The atlas
Eight acts, a coda, three appendices
The atlas proper is one continuous instrument, read by scrolling: The Instrument, One Hand, The Law, The Tariff, The Encounter, The Higher Court, The Measure, Eras & Territories, a coda in the archive, and three appendices — thirty-three plates in all. Each plate asks one question, answers it with one figure, and states in its own caption what it cannot answer.
The same rail that opens the atlas, condensed: the wedge widens as the frame of reference does, from one sentence to the whole nine-year record. Each row opens its act.
Three habits make it read faster. Every technical term is defined on hover. Every number opens: click into a figure and it will keep going down until a real reviewer sentence, with a link to the OpenReview thread it came from, is on screen. And every plate has a machine-readable deposition — the same claim in JSON, with its counts, at /api/v1/plates/.
Some of what the plates hold
- 71.8%of all units resolve negative — the memo layer articulates the logic of criticism far more often than praise (a property of the instrument as much as of reviewers).
- 138/144the grammar: each object of scrutiny has a canonical standard — novelty is judged by the novelty standard, cost by cost-benefit — and the coupling barely moved in nine years: 138 of 144 cells sit inside the shuffle-null floor.
- 26→6→13%the script: reviews open conceptually (framing 1.9× base rate) and close with hygiene (reproducibility 1.4×); praise collapses mid-review and returns at the close.
- 4.6%the rebuttal: of 130,650 post-response units, only 5.6% weaken or reverse (4.6% weakened, 1.0% reversed); even fresh experiments strengthen the original judgment ~3× more often than they soften it.
- −10.6ppthe tribunal: novelty criticism shows the largest acceptance gap (association only, not causation); compute-cost and robustness criticism show none at all.
- 0.65the legible verdict: counting what was criticised reads a paper’s accept-or-reject at AUC 0.65 (its words 0.74; the panel’s own scores 0.90) — and that legibility has fallen for nine years. (The old “gauntlet” slope — acceptance falling with each added front of criticism — turned out to be expected-by-construction and now hangs in the Null Cabinet.)
- 38.6%dissent: reviewers of the same paper split most on problem framing — whether the question matters — and least on robustness, where evidence speaks.
- −13.7ppthe kinder judge: meta-reviewers run 13.7 points less negative and 10.5 points more positive than the panels they summarize.
- +3.2ppthe drift: 2018→2026, scrutiny moved from theory and clarity toward compute cost, statistical rigor, and reproducibility — from what is this idea toward what does the evidence cost.
- 7.8→3.1%the watermark: a frozen thirteen-word LLM-vocabulary signature appears in 7.8% of 2024 reviews (the pre-2023 trend expects 1.6%), then fades to 3.1% by 2026 — while co-reviews of the same paper grow 14% more alike under a fixed vocabulary. The fade measures the practice’s visibility, not its prevalence; no single review is accused.
Every figure above is interactive in the atlas, with counts, definitions, and raw specimen units behind it. None of them is the point of the page you are reading; they are here so that you know what kind of thing is on the other side of the door.
The open door
Nothing here is behind anything
The corpus was public before this project touched it, the derived data is public now, and so is every step in between. There is no private supplement, no held-back table, no "results available on request".
The circuit of the open door. The record was public before the project began; the code that reads it and the data it becomes are public now; and the two readers — the atlas for people, the API for machines — stand on the same 50 islands, with the method walking beside every arrow. Each node above is a door: it links to the public place it names.
llms.txt and an OpenAPI schema.The ledger
What it cost to build, in full. The bill is small enough to be worth stating exactly, because the point of stating it is that anyone could pay it.
| Stage | Scale | Cost (USD) |
|---|---|---|
| Collection + normalization (OpenReview) | 52,460 forums · 792,703 messages | 0 |
| DeepSeek memos — Full Layered (2026) | 151,193 memos | 314.42 |
| DeepSeek memos — Direct (2018–2026) | 51,813 memos | 194.75 |
| Qwen structuring — both tracks, incl. retries | ~1.42M units emitted | 103.59 |
| Taxonomy, embeddings, aggregation, atlas | local compute only | 0 |
Pilot-phase exploration (episode schemas, atlas cards, calibration runs) is additional and itemized in the project's provenance ledger; the figures above are the production spine.
Reproduce it
The companion page, The Atlas Method, records the full pipeline — every script, model, parameter, seed, cost, failure mode, and the exact commands — at the level of detail needed to rerun it, along with an honest account of what cannot be reproduced (LLM sampling nondeterminism, a handful of one-off manual repairs, and upstream drift in OpenReview itself).