Atlas of Judgment · the companion

What this is

theatrum judicii — the theatre of judgment
rise through the tiersnine tiers, and a colophon
Tier I

The project

theatrum — the room this is modelled on
I · what it is

An attempted description of peer review, at the resolution of a single thought

The Atlas of Judgment is a metascience instrument: an analysis of every public peer review of ICLR — the International Conference on Learning Representations — asking what the act of reviewing is. To make that analysable, each review is read into atomic units of evaluative logic — and what kinds of unit the record turns out to contain is itself part of the question — then the units are counted, compared, and tested at scale. Not what did reviewers decide, but how did they think.

Peer review decides what counts as knowledge in machine learning. It runs on the unpaid attention of tens of thousands of people, it is argued about constantly, and its interior is studied far less than its verdicts. What is visible from outside is the output — a score, an accept, a reject — the way a verdict is visible while the deliberation behind it is not. This project is an attempt to open the deliberation and describe it.

The result is 1,420,178 atomic units of reviewer reasoning, drawn from nine years of ICLR and analysed across thirty-three plates. It is not an argument about peer review. It is an attempt at a description of it — made in one chosen form, carried all the way through — and it was built so that any part of the description can be taken apart by someone who doubts it.

The whole instrument, at a glance
conspectus totius operis
The record
Nine years of ICLR, public on OpenReview — every review, rebuttal, and verdict, rejected papers included.
52,460 submissions
792,703 messages
dissected
The unit
Each criticism cut into one movement of a reviewer’s mind: inspected · observed · reasoned · judged.
1,420,178 units
classified
The taxonomy
Every unit assigned in a grammar induced from the data: 12 objects of scrutiny × 12 reasoning standards.
144 cells
counted, drawn
The atlas
Every figure a count of units — and every number a door, opening back down to a reviewer’s sentence.
33 plates · 8 acts

The whole project in one strip: the public record, the four-part unit it is dissected into, the 12 × 12 grammar each unit is filed under, and the plates that count them. Read backwards, it is the route of doubt: any number on any plate descends, step by step, to the reviewer sentences beneath it.

410,586 units (ICLR 2026, review-level) · 1,009,592 units (2018–2026, forum-level) · 52,460 submissions collected · verified 2026-08-18
Tier II

The motive

cur ratio, non numerus — why the reasoning, not the number
II · why the reasoning, and not the score

A score is a judgment with the reasons thrown away

Two reviewers write 5. One means the idea is not new. The other means the experiments do not support the claim. Averaged into a panel statistic, the difference disappears — and the difference was the whole of what happened. A score is a compression, and what it compresses away is precisely the part that could be argued with, learned from, or improved.

The project's research question was fixed early, and verbatim, in its design documents: not whether a paper was accepted, nor whether a review was favorable, but — what do reviewers inspect, and through which criteria, comparisons, assumptions, and inferences does that observation become a judgment?

Peer review is a central quality-control mechanism of science, yet the reasoning inside it is rarely studied as an object in its own right. Reviews are usually reduced to scores, sentiment, or accept/reject outcomes. This project treats the reasoning itself as the data.

That choice sets the shape of everything downstream. Sentiment analysis would have counted how reviewers felt. Outcome modelling would have predicted what became of the paper. Neither says what was being weighed. The atlas is built for a narrower and more stubborn question: which part of a paper was looked at, and which standard was held against it. Asking that at scale meant choosing a grain and a vocabulary before any answer could exist — the four-part unit, the twelve-by-twelve grammar — so every answer the atlas gives is an answer in that chosen vocabulary: one way of carving the reasoning, carried through, not the only way it could have been carved.

Tier III

The venue

unum forum — one forum, and only one
III · why ICLR

The venue that publishes the reviews of the papers it rejects

One property makes ICLR nearly unique among top machine-learning venues: it publishes the reviews of rejected papers. Public review archives usually begin where acceptance ends — and a record without the rejected majority shows judgment only where judgment said yes. ICLR's OpenReview forums carry the full distribution — every submission, every review, every rebuttal, every verdict — for nine consecutive years, 2018 through 2026. That is 52,460 submissions and 792,703 forum messages, all from a single source with a single collection path.

One forum, as anyone may read it
acta integra — the record, entire
The submission
the paper itself, public from the day the forum opens — and it stays.
The official reviews ≈ 3.8 per paper
each on one form — summary · strengths · weaknesses · questions — with a rating and a confidence attached: ≈425 words of judgment, signed by a stable pseudonym.
The exchange
the authors answer under every review, and reviewers may return — the whole volley, timestamped, stays in the thread.
The meta-review
the area chair reads the panel and writes reasons of their own.
The decision
accepted · record stays open rejected · record stays open
52,460 submissions · 792,703 public messages · nine years, one source

One forum’s record, drawn to its parts. The counts are the corpus’s own: ≈3.8 official reviews per paper and ≈425 words per review are the 2026 means (74,380 reviews over 19,474 papers); 52,460 submissions and 792,703 messages are the nine-year totals. The fork’s two ends are drawn with the same ink on purpose — the record’s completeness does not depend on the verdict, and that is the property this atlas rests on.

The choice is a constraint, not an endorsement. ICLR is one venue, in one field, with one culture and a review form that changed four times across the nine years. Everything in the atlas describes ICLR; nothing in it should be read as a fact about peer review in general unless the plate says so, and none of them do.

Tier IV

The atom

unitas judicii — one unit of judgment
IV · the unit of analysis

One complete movement of a reviewer's mind

The unit of analysis is what this project treats as one complete movement of a reviewer's mind — from a thing inspected, through an observation, through the standard invoked, to a judgment. A real unit from the corpus:

Inspected
The paper's notation system and formatting consistency.
→ Observed
The reviewer finds the notation "difficult to follow" due to inconsistencies in subscript/superscript usage.
→ Reasoned
Clear notation is a prerequisite for verifying correctness; ambiguity here creates a barrier to evaluating technical soundness.
→ Judged
The presentation quality is insufficiently rigorous — a formal weakness that hinders scrutiny.

Specimen № 335 of 410,586, from ICLR 2026 submission hPUjTTj64O. The same unit is opened part-by-part on the atlas's own Plate 0.

Each unit also carries a verdict polarity (negative / positive / conditional / uncertain / mixed), a suggested fix when the reviewer offered one, line-level citations back to the source review, and a flag recording whether it is grounded in the reviewer's explicit words or inferred by the analysis layer.

Everything in the atlas is a count of these. No plate holds anything that is not, at bottom, a pile of units — which is why every figure can be opened until a sentence written by an actual reviewer is showing. A number here is a door, not a conclusion.

Tier V

The making

manus artificis — the maker's hand, shown
V · how it was made

Three layers, kept separate on purpose

No one reads 792,703 forum messages. The scale that makes the question answerable is exactly the scale that puts a machine between the reader and the reviewer. The honest way to handle that is not to hide the machine but to name the layers and keep them apart, so that a claim can always be traced to the layer that produced it.

Layer IHuman reviews — the raw OpenReview record, collected once and kept immutable. No anonymization is applied; the data is already public.
Layer IIAnalytic memos — a DeepSeek model reads each review (or forum) and writes a free-text metascientific analysis, under prompts that forbid re-reviewing the paper, require line-ID citations, and deliberately hide outcomes and other reviews at the initial stage.
Layer IIIStructured logic units — a Qwen model normalizes each memo into schema-constrained JSON. It never sees the raw review; it is a normalization layer over Layer II, not a second opinion.

Two tracks, one taxonomy

Two independent extraction tracks cover complementary views. The 2026 track works review-by-review at high resolution: 74,380 reviews → 410,586 units. The 2018–2026 track works forum-by-forum across nine years: 50,861 forums → 1,009,592 units, with temporal fields (before/after rebuttal, judgment change) that power the drift and rebuttal analyses.

Both tracks share one taxonomy — 12 objects of scrutiny × 12 reasoning standards — induced from the data itself: a 12,000-unit sample was embedded and clustered with no seeding from prior schemes, the clusters were read and named by hand, and every unit in both tracks was then assigned to its nearest category. Definitions for every term are in the atlas lexicon, and on hover throughout.

Tier VI

The limits

quod non est — what it is not
VI · what this is not

A machine reading of reviews, not a measurement of reviewers

What this is not: a direct measurement of human reviewers. Every unit is a two-stage machine reading. The support_status flag (explicit vs. inferred) and a per-category reliability grid in the atlas appendix exist so that every claim can be read against the instrument's own error.

Read the atlas with five caveats. (1) It measures a machine reading of reviews, not reviewers directly; the reliability grid shows where inference substitutes for quotation. (2) The negative share reflects the memo prompts' bias toward articulating criticism — it is not a sentiment ratio. (3) Coverage is 98.1–98.2% on both tracks; the missing tail (parse failures, provider content-policy refusals) is documented and its topic correlation unverified. (4) Decision-linked plates (the Consequence) are associations: outcomes were never shown to the extraction pipeline, but criticism breadth and paper quality are entangled by design. (5) The four-part unit and the 12 × 12 taxonomy are one induced carving of the text, not the only possible one: their internal reliability is measured, and the categories track ICLR's own sub-scores where they should, but no check can establish them as the right carving — a different scheme would draw a different map of the same reviews.

A caveat kept in an appendix is a caveat that has been hidden. These ride at the same weight as the findings, on the plates themselves as well as here — and the atlas's third appendix is given over entirely to results that failed: patterns that looked real, were tested against a null, and did not survive it.

Tier VII

The atlas

tabulae XXXIII — thirty-three plates
VII · how to read it

Eight acts, a coda, three appendices

The atlas proper is one continuous instrument, read by scrolling: The Instrument, One Hand, The Law, The Tariff, The Encounter, The Higher Court, The Measure, Eras & Territories, a coda in the archive, and three appendices — thirty-three plates in all. Each plate asks one question, answers it with one figure, and states in its own caption what it cannot answer.

The same rail that opens the atlas, condensed: the wedge widens as the frame of reference does, from one sentence to the whole nine-year record. Each row opens its act.

Three habits make it read faster. Every technical term is defined on hover. Every number opens: click into a figure and it will keep going down until a real reviewer sentence, with a link to the OpenReview thread it came from, is on screen. And every plate has a machine-readable deposition — the same claim in JSON, with its counts, at /api/v1/plates/.

Some of what the plates hold

  • 71.8%of all units resolve negative — the memo layer articulates the logic of criticism far more often than praise (a property of the instrument as much as of reviewers).
  • 138/144the grammar: each object of scrutiny has a canonical standard — novelty is judged by the novelty standard, cost by cost-benefit — and the coupling barely moved in nine years: 138 of 144 cells sit inside the shuffle-null floor.
  • 26→6→13%the script: reviews open conceptually (framing 1.9× base rate) and close with hygiene (reproducibility 1.4×); praise collapses mid-review and returns at the close.
  • 4.6%the rebuttal: of 130,650 post-response units, only 5.6% weaken or reverse (4.6% weakened, 1.0% reversed); even fresh experiments strengthen the original judgment ~3× more often than they soften it.
  • −10.6ppthe tribunal: novelty criticism shows the largest acceptance gap (association only, not causation); compute-cost and robustness criticism show none at all.
  • 0.65the legible verdict: counting what was criticised reads a paper’s accept-or-reject at AUC 0.65 (its words 0.74; the panel’s own scores 0.90) — and that legibility has fallen for nine years. (The old “gauntlet” slope — acceptance falling with each added front of criticism — turned out to be expected-by-construction and now hangs in the Null Cabinet.)
  • 38.6%dissent: reviewers of the same paper split most on problem framing — whether the question matters — and least on robustness, where evidence speaks.
  • −13.7ppthe kinder judge: meta-reviewers run 13.7 points less negative and 10.5 points more positive than the panels they summarize.
  • +3.2ppthe drift: 2018→2026, scrutiny moved from theory and clarity toward compute cost, statistical rigor, and reproducibility — from what is this idea toward what does the evidence cost.
  • 7.8→3.1%the watermark: a frozen thirteen-word LLM-vocabulary signature appears in 7.8% of 2024 reviews (the pre-2023 trend expects 1.6%), then fades to 3.1% by 2026 — while co-reviews of the same paper grow 14% more alike under a fixed vocabulary. The fade measures the practice’s visibility, not its prevalence; no single review is accused.

Every figure above is interactive in the atlas, with counts, definitions, and raw specimen units behind it. None of them is the point of the page you are reading; they are here so that you know what kind of thing is on the other side of the door.

Tier VIII

The open door

janua aperta — everything that is public
VIII · everything that is public

Nothing here is behind anything

The corpus was public before this project touched it, the derived data is public now, and so is every step in between. There is no private supplement, no held-back table, no "results available on request".

The circuit of the open door. The record was public before the project began; the code that reads it and the data it becomes are public now; and the two readers — the atlas for people, the API for machines — stand on the same 50 islands, with the method walking beside every arrow. Each node above is a door: it links to the public place it names.

The ledger

What it cost to build, in full. The bill is small enough to be worth stating exactly, because the point of stating it is that anyone could pay it.

StageScaleCost (USD)
Collection + normalization (OpenReview)52,460 forums · 792,703 messages0
DeepSeek memos — Full Layered (2026)151,193 memos314.42
DeepSeek memos — Direct (2018–2026)51,813 memos194.75
Qwen structuring — both tracks, incl. retries~1.42M units emitted103.59
Taxonomy, embeddings, aggregation, atlaslocal compute only0

Pilot-phase exploration (episode schemas, atlas cards, calibration runs) is additional and itemized in the project's provenance ledger; the figures above are the production spine.

Reproduce it

The companion page, The Atlas Method, records the full pipeline — every script, model, parameter, seed, cost, failure mode, and the exact commands — at the level of detail needed to rerun it, along with an honest account of what cannot be reproduced (LLM sampling nondeterminism, a handful of one-off manual repairs, and upstream drift in OpenReview itself).

Tier IX

The colophon

sub finem — at the end, the terms