Atlas of Judgment

How a Paper
is Judged

In 1522, anatomists moved the name atlas from the shoulders to the first vertebra — the joint that lets a head nod yes, or turn away. This one is engraved with 40,000 real units of reviewer reasoning, and it moves as they conclude: mostly, it turns away.

the atlas vertebra · c1 — nod: assent, turn: refusal · engraved with 40,000 real units of reviewer reasoning · touch the bone to read them · light of 2018
Descend
the question this atlas asks

What do peer reviewers actually look at — and how do they decide?

To find out, open-weight language models read every public ICLR review and split each one into pieces of reasoning. One piece — a unit — records four things: what the reviewer inspected, what they observed, the standard they judged it by, and what they concluded. This page counts 410,586 units from ICLR 2026 (98% of its reviews), read one review at a time, and 1,009,592 units from 2018–2026, read a second way that takes in each paper's whole discussion — rebuttal replies and the area chair included. Every number below is a count of units. The reading is a machine's, not a human's; the note below explains what that costs, and every plate links back to the original reviews.

410,586
logic units
74,380
reviews
19,474
papers
+ 1,009,592 units spanning 2018–2026 · taxonomy induced, not imposed · ICLR MMXXVI
the sky behind this page drifts in the corpus's true verdict proportions
Plate 0 · one unit, opened — hover or tap a part to find it in the sentencespecimen № 335 of 410,586
① inspected
the novelty of the method’s early-commitment mechanism — the part of the paper this judgment is about
③ standard
the norm the reviewer holds: novelty requires a new mechanism, not just a re-packaging of known search strategiesits leader ends in the air: this is written nowhere in the sentence. Most standards are implicit, and the machine’s reading names them.
“Limited novelty: Core idea — early-commitment to promising partial solutions — closely resembles existing beam/pruning and adaptive-consistency techniques; incremental contribution.”
② observed
the core idea resembles existing beam/pruning and adaptive-consistency techniques — what the reviewer saw there
④ verdict
the contribution is incremental negative
Unitas judicii, quattuor partibus apertaone weakness from an official review of ICLR 2026 submission hPUjTTj64O ↗ · matched verbatim to the raw OpenReview record, 2026-08-25

Four parts, one unit. This exact dissection was performed 410,586 times on ICLR 2026, and a million more across nine years of discussion — every figure below is a count of these. The note that follows says what a machine’s reading costs, why this particular cut is a choice, and how to check both.

Before you start · what was counted, and how
A unit is one piece of a reviewer's reasoning

Open-weight language models read each review and split it into the separate judgments it makes. One unit records four things: the part of the paper inspected, what was observed there, the standard the reviewer judged it by, and the verdict reached. A typical review yields five or six. Every figure on this page is a count of units, so the labels are a machine's reading of the text — consistent across a million reviews, but not ground truth. The grain is a choice, too: these four parts, and the 12 × 12 taxonomy the plates sort them into, are one induced way of carving reasoning, held fixed throughout — internally consistent by the Method page's checks, but not a carving whose aptness any plate can test. A different scheme would draw a different map. Units link back to the review they came from; check them.

Two readings, not one corpus

The reviews were read twice, in two different ways, and the plates use whichever fits the question. ICLR 2026 was read one review at a time (410,586 units), so each unit can point at the line of the review it came from. 2018–2026 was read a second way — each paper's whole discussion in a single pass, including the replies written after the author's rebuttal and the area chair's meta-review (1,009,592 units) — which is what makes it possible to see when a judgment was made and whether it later changed. Neither is a subset of the other. Each plate's header names the years it uses; the Method page gives both pipelines in full. One more thing worth knowing before the sequence starts: the order of the plates is an editorial arrangement, made after the analyses existed, not the order in which the work was done — the Method page carries a plate-by-plate ledger of which questions were written down before the numbers existed, which plates are frankly exploratory, which we simply have no record for, and one claim that was published wrong and corrected in place. Finally, the stage-setting plates travel as overtures at the heads of the acts they serve: the calendar of one full review season opens Act V, one warning about review timing — reviews filed weeks early run harsher and longer than deadline-day ones, and no plate controls for filing time — opens Act VII as the Tide, and the venue's own census opens Act VIII.

The whole instrument, and the order it is shown
conspectus et cursus operis
The record
Nine years of ICLR, public on OpenReview — every review, rebuttal, and verdict, rejected papers included.
74,380 reviews · 19,474 papers
in 2026 alone
dissected
The unit
Each criticism cut into four parts — one chosen grain, held fixed: inspected · observed · standard · verdict.
410,586 + 1,009,592 units
classified
The taxonomy
Every unit filed in a grammar induced from the data: 12 objects of scrutiny × 12 reasoning standards.
144 cells
counted, drawn
The atlas
Every figure below is a count of these units — and every number a door, opening back down to a reviewer’s sentence.
plates I–XXX · eight acts & a coda
arranged
the course · each act one step farther from the sentence
ACT I
The Instrument
how a million reviews became countable units · plates I–III
Plate I · The Anatomy
Corpus: ICLR 2026 · 410,586 units from 74,380 reviews

Object → standard → verdict

Why start here
the question this project began with — recorded before any data existed — was: what do reviewers actually look at, how do they think, and how do they decide? This diagram is that question answered at its coarsest grain, and it is the closest thing the atlas has to a portrait of its subject: every act of judgment taken apart into what was inspected, the standard it passed through, and the verdict where it landed — all 410,586 units of ICLR 2026 at once. Every later plate is a slice of this picture. Each unit of reasoning was labelled three ways: the object it inspects (what part of the paper the reviewer was looking at — 12 of them), the standard it argues from (the rule the reviewer used to turn that observation into a judgment — 12 of them), and the verdict it reaches (negative, positive, mixed, conditional, or uncertain). One real unit, walked through its three labels: “A novel contribution supported by thorough empirical evaluation warrants high marks, even if some clarity issues exist” — its object is novelty (that is what was under inspection), its standard is merit recognition (the rule that strengths are credited to the score), and its verdict is positive. Every ribbon below is thousands of such triples stacked. Both label sets were induced from the reviews themselves, not chosen in advance. The couplings are the content: ribbon thickness is the strength of each bond — how many units tie an object to a standard, and a standard to a verdict — so a thick ribbon is a habit of the discipline and a thin one an exception. Hover a node to trace everything its units touch, upstream and down, and click one to read real units from it; hover any term for its definition, or see the full list.
71.8% of everything reviewers write is criticism — and two thirds of the praise reaches its verdict through a single standard, merit recognition.
BAR HEIGHT = SHARE OF ALL 410,586 UNITS
Fig. 1a
Result
five things the ribbons show. First, criticism dominates every path: 71.8% of all units end negative, and no object escapes — the gentlest-treated object still runs 60% negative, the harshest 88%. That is not the acceptance rate in disguise: even reviews awarding an 8 are made of 45% negative units — thorough praise is still built out of criticism. Second, praise has a bottleneck: — mercy is not spread across the reasoning; it has its own channel. Third, the single largest current in the corpus is (4.8% of everything): the modal act of ICLR reviewing is asking why the experiments were set up as they were. Fourth, the bonds themselves run from tight to loose: — one standard (its own) carries 67% of its units and two cover 80% — while statistical rigor is the most promiscuous, spreading so widely that its top standard holds only 14% and eight are needed to cover 80%. The median object needs six. Judging what a paper is has one law; judging whether its numbers hold has many. Fifth, the strongest bonds that are not an object paired with its own namesake standard — the ones that carry information: , so the citation list is read chiefly as a claim about priority rather than a bibliography to complete; ; and — the motivation section is the field's default site of compliment.
Reading it
one caution on the 71.8% level itself: the models that split reviews into units catch explicit criticism more reliably than brief praise — a passing "well written" often goes unrecorded — so the negative share is tilted upward, and it measures the reading as much as the reviewers (Appendix II reports the check). Comparisons stay safe, because every object and standard was read the same way. How each object × standard pairing resolves, cell by cell, is the map below.
Fig. 1b — The severity map: where mercy lives
All outside the merit-recognition column condemn more often than not; inside that one column, criticism falls to 14–35% for eleven of the twelve objects.
Fig. 1b
Why this is here
this is not new data — it is the diagram above shown a second way. The sankey holds the verdicts too, but reading how a particular pairing tends to end means hovering its ribbons one at a time; here the same 144 pairings are laid out so the outcome of every one is visible at once and can be compared at a glance.
How to read
the same 144 object × standard pairings, as a grid: rows are objects, columns are standards, and each cell is coloured by the share of that pairing's units that end in a negative verdict. So a deep red cell means: when a reviewer examines this part of a paper under this rule, the judgment almost always ends in criticism; a pale cell means that pairing tends to end well for the author. Hover any cell for its exact share and unit count; one cell of the 144 (novelty × the reproducibility norm, 36 units) is too thin to read and is hatched grey rather than coloured.
Result
the map has one exception, not many. , every single one of the 131 pairings with enough units runs majority-negative — the mildest of them () still condemns 56% of the time. the share falls to 14–35% for eleven of the twelve objects. , named as the extremes they are (each holds hundreds to thousands of units, so the shares are not sampling noise): related work judged by presentation-as-trust (95.9% negative), clarity judged by the same standard (94.7%), and reproducibility judged by statistical identifiability (94.4%). And one object finds no mercy anywhere: related work occupies the harshest cell of the whole map and is the worst-treated object inside the mercy column itself — , where no other object exceeds 35%. That squares with Fig. 1a's fifth result: a citation list is read as a claim about priority, and there is no standard under which that claim is usually granted.
Reading it
the wall of red is also a reminder that the pipeline records explicit criticism more readily than passing praise — the level is tilted; the comparisons between cells are what to trust.
Fig. 1c — The instrument, checked against the venue’s own gauges
ICLR asks its reviewers for three sub-scores of their own. Where this atlas’s categories and those gauges should agree, they do — — and the one place they part is itself a finding.
CELL = SUB-SCORE GAP, CRITICIZED VS NOT, AT FIXED OVERALL RATING · 1–4 SCALE · ICLR 2026
Fig. 1c
Why this is here
every plate that follows leans on these twelve categories. Their induction is documented in the method pages; this figure is the check a reader can hold in one glance — the taxonomy against a rubric it never saw, collected by the venue itself.
How to read
rows are the twelve objects of scrutiny; columns are ICLR’s own three sub-scores (soundness, presentation, contribution, each rated 1–4 by the same reviewer). A cell is the gap in mean sub-score between reviews that criticise that object and reviews that do not, measured inside each overall-rating level and then averaged — so a review’s general harshness is held fixed, and what remains is where the reviewer’s own gauges dip. If the twelve labels were arbitrary, no gauge would care which object was criticised: every cell would sit near zero. Rows are ordered by their deepest dip.
Result
agreement where agreement is owed: clarity criticism dents presentation (−0.35, the sharpest cell of the map, four times any other), novelty criticism dents contribution (−0.12, that column’s deepest), and the evidence dockets — reproducibility, statistics, baselines — dent soundness most. Two smaller marks are worth the detour. Theory criticism dents presentation three times harder than soundness (−0.07 vs −0.02): “I could not follow the theory” is filed, by reviewer or by reading, as a writing complaint rather than a rigor complaint. And a clarity charge sits slightly higher on contribution (+0.04): at the same rating, a paper faulted for its prose keeps its ideas’ credit.
Reading it
sub-scores exist only in recent forms, so this check runs on ICLR 2026 alone (74,380 reviews); it is association at fixed rating, not causation; and the gauge cannot split the reviewer from the instrument — “theory filed as writing” may be the reviewer’s habit, the model’s reading, or both. What it does establish is that the categories carry the venue’s own distinctions, not just the model’s.
Plate II · The Grammar
Corpus: ICLR 2026 · 410,586 units · Fig. 2b’s movement: 2018–2026

The law is often borrowed

The question
Plate I showed which objects reviewers inspect and which standards they argue from. Are the two coupled? For each object: how is its criticism argued?
What it bears on
if any standard could be applied to any part of a paper, the standard label would carry no information — which would in turn empty out the later plates that compare reviewers by the rule they invoked. That consequence is a reason to check the coupling, not the reason this plate was built; it is one of the first things we drew. Every row is one object, divided by the standards used on it. A deep diagonal is the discipline judging each thing by its own rule; the off-diagonal cells are where one subject is judged by another subject's standard. One motive for showing nine years of movement this early, in Fig. 2b: if these couplings drifted freely year to year, the plates that pool nine years of data later would rest on sand — 2b is the stability check that licenses them.
Five of twelve objects answer chiefly to their own law; the rest are judged under borrowed ones — design justification above all.
Fig. 2a
How to read
this grid is not a new dataset either — it is a third view of the same 144 object × standard pairings as Plate I: the sankey drew them as ribbons, the severity map coloured each by how it ends, and this one colours each by how common it is within its row. Frequency is exactly what this plate's question — are object and standard coupled? — needs: if any standard could be applied to any object indifferently, every row of this map would be the same stripe pattern — each column simply at that standard's corpus-wide share — and knowing the object would tell you nothing about the rule. Coupling is the departure from that: rows with individual profiles, each object concentrating its units on its own few standards. That departure is what the map shows everywhere — no two rows match. Each row is one object of scrutiny; cells are the share of that object's units argued under each standard, row-normalized, deeper brass for more. The one cell too thin to read (novelty × the reproducibility norm, 36 units) is hatched grey here as everywhere. If the object/standard distinction feels slippery here — what separates novelty the object from the novelty standard? — two real units about the same object show it. The object is what the reviewer was looking at; the standard is the rule the argument runs on. “Lack of sufficient differentiation from prior art undermines contribution” looks at novelty and argues from the novelty standard — the rule that being distinct from prior work is itself the requirement. “A novel contribution supported by thorough empirical evaluation warrants high marks, even if some clarity issues exist” looks at the same object but argues from merit recognition — the rule that strengths should be credited to the score. Same subject, two different laws; this plate counts how often each law is the one applied.
Result
counting each row's most-used law: five objects are judged first by their own named standard — novelty (67%), clarity (48%), robustness (33%), compute cost (32%), reproducibility (31%) — while ; the citation list answers first to the novelty standard (50%), and problem framing to merit recognition (29%). Even the tightest coupling is far from total: novelty argues a third of its units under other laws, and — the standard label carries information the object's name does not.
Reading it
treat diagonal brightness as bookkeeping, not certainty — several standards are name-twins of their objects (novelty and the novelty standard, robustness and the robustness norm), and for those cells a deep diagonal is partly analytic, since both labels can be responses to the same sentence; the evidence lives in the off-diagonal mass — statistical rigor judged by identifiability, empirical scope answering to the claim-evidence matching norm — and in how the whole grammar moved (Fig. 2b).
Fig. 2b — The grammar's nine years: where P(standard | object) itself moved
The grammar is a habit, not a fashion: — and what little moved was mostly ground the novelty standard gave up.
THE FIRST OF THREE NINE-YEAR VERDICTS — this map says the coupling grammar barely moved. Act VIII returns to the same question twice more: what reviewers attend to did move (Plate XXVI), and what the stated law of novelty says did not (Plate XXVII).
Fig. 2b
Why it is here
Fig. 2a is a still photograph — the coupling as it stands now. The plate's question has a natural second half: is that grammar a fixed habit of the discipline, or has it been drifting? This figure answers with the same matrix, except each cell now shows the change in P(standard | object) from the earliest era the corpus covers (2018–19, 32,257 units) to the current one (2025–26, 606,120 units).

What chance would draw
if the grammar were fixed, the two eras would differ only by sampling noise — so the noise is measured rather than assumed: shuffle the era labels within each object and redraw the map, and in 95% of 50 relabelings even the single largest cell change stays within ±4.05 percentage points. Cells inside that floor are dimmed to near-invisibility; under the null, this entire map would fade to the unmarked page.

How to read
brass = the standard gained jurisdiction over that object, violet = it lost it; deeper is a larger move; hover any cell for its early and current shares. (Those shares come from the nine-year labeling run, whose levels differ somewhat from Fig. 2a's 2026 corpus — this map's subject is the change within one consistent pipeline, not the levels.)

Result
138 of the 144 cells stay inside the shuffle floor: across nine years in which the corpus grew almost nineteen-fold, the grammar mostly did not move. Six cells did. Two are specialization — and . Three sit in a single row: . And .

Is the movement substantive?
A share can move for boring reasons — reviews now arrive in structured Strengths/Weaknesses templates, and framing units are indeed praise more often than they were (30% → 39% positive), which by itself inflates merit recognition, the crediting rule. So each mover was re-checked within criticism and praise separately. Merit recognition's gain does shrink under that control but survives (+6.3pp among praise units alone); every other mover passes untouched — the novelty standard recedes among critical and praising units alike (−4.5pp and −5.4pp on problem framing, −4.3pp and −7.7pp on related work), and cost-benefit gains most inside critical units (+10.3pp), where a mere change of tone could not put it. Reading the units confirms that the rules themselves did not change meaning: a 2018 cost unit argues “fair comparison requires accounting for all computational resources consumed” and a 2025 one “overhead claims must be backed by runtime experiments” — the same law, reached for more often. What these checks cannot exclude is a quieter drift in phrasing that the labeling model reads as a different rule; every level here rests on one model's reading, applied uniformly to both eras.

Interpretation
after the checks, the drift sorts into three grades of confidence. Firmest: the quiet specializations — cost and reproducibility questions each acquired their own law — and the novelty standard's retreat from the two objects next door to it, framing and citations, which holds among critical and praising units alike. To be read more gently: the swing toward merit recognition, the one mover that partly reflects the era's style of reviewing — praise now arriving in its own labelled section — and only partly a change in which rule judges framing. And the widest finding is the stillest one: 138 cells, nine years, a corpus grown nineteen-fold, and no movement that shuffling the era labels could not have drawn.
Plate III · The Rhetoric
Corpus: ICLR 2018–2026 · 1,009,592 units

Each standard has a voice

The question
the two plates before this recorded what a reviewer judged and which rule they judged by. Neither says anything about the shape of the argument — whether the reviewer asserted a norm, reported that a check could not be made, offered a competing explanation, or weighed the evidence the paper offered. Two criticisms can share an object and a standard and still be different acts. So this plate classifies the argument itself: not just which standard fires, but the shape the reviewer's reasoning takes. Six forms, assigned to all 1,009,592 units of the nine-year corpus, 2018–2026 (the second of the two readings described above: whole discussions, not single reviews) by a locally-trained classifier (no API, no keywords-only shortcut). The six are defined below the figure title, each with a real unit from the corpus. Where the six come from matters: they are induced from examples, not written in advance. A first pass noticed recurring argument language in real units — without X…, should…, could be an artifact…, compared to prior work…, beyond this setting… — and used those surface markers to draw a stratified sample. The analyst then hand-read 600 of its units and wrote the codebook against the text (eight codes, merged to six), overruling the markers freely: within marker-flagged units the hand code agreed with the marker’s suggestion only 61% of the time, and the corpus’s largest form — on-balance weighing — has no marker at all; it emerged from the unflagged residue. A classifier scales that reading to the full corpus. One defensible carving of the space, not the only one; the caption owns the method and its error.
Fig. 3 — The voice of each standard — six argument forms, all 1,009,592 units
Five of the six argument forms almost always criticize. When praise comes, nearly four times out of five (78.8%) it arrives as on-balance weighing.
NORM — the reviewer asserts a bar the paper must clear. “Clear problem framing is essential for evaluating contribution.”
CAN’T-VERIFY — a check the reviewer needs cannot be made from what the paper provides. “Without variance metrics, limited performance improvements cannot be assessed for significance.”
PRECEDENT — prior work is held against the paper. “If the core contributions are already present in cited prior work, the novelty claim is weakened.”
SCOPE — the claim is asked how far it carries. “Strong assumptions limit the method’s generalizability to real-world data with feature correlations.”
RIVAL-CAUSE — another explanation could produce the same result. “Performance differences may reflect training budget artifacts rather than method ceilings.”
ON-BALANCE — the evidence offered is weighed and found sufficient or wanting. “Improved clarity increases value, but broader experimental validation is needed to confirm generalizability.”
EACH BAR = THAT STANDARD’S UNITS, SPLIT BY ARGUMENT FORM
Fig. 3
How to read
each row is one reasoning standard; its bar splits that standard’s units by the argument form the classifier assigned, and the right column names the dominant form. The plate’s question — does each standard have a voice? — is answered by how far the rows depart from sameness: if the shape of an argument had nothing to do with the rule it runs on, every bar here would show the same six-way split, the corpus-wide mix (37% on-balance, 30% norm, 13% scope, 8% precedent, 6% rival-cause, 6% can’t-verify) — drawn as the dimmed first bar, so each row's departure from it is the finding.
Result
the rows are nothing like the uniform mix. ; — more than seven times that form’s corpus-wide share — the form that says a check the reviewer needs cannot be made, as in the 2019 unit “Without error bars, it is impossible to assess if differences are statistically significant”; and (29% of its units — though even there, on-balance weighing leads the row). On-balance is also where the mercy lives: (42% positive, and 78.8% of all praise in the corpus takes this one form; the other five forms run 88–97% negative).
Method
600 units labeled by the analyst (an LLM reading each sentence), merged to six forms; the example quoted beside each form above is one of those 600 gold units, verbatim. A logistic classifier over local embeddings plus surface-marker features (5-fold CV accuracy 0.685, macro-F1 0.666; the weakest class is the precedent-citing form, at F1 0.53 — read its cells gently) then labels the full corpus. A cheaper second annotator agreed with the gold labels only 42% of the time and was excluded — the construct is genuinely interpretive, and these are one consistent reading at scale, not ground truth.
Reading it
the six-way split is one model's reading of a genuinely interpretive construct (macro-F1 0.666, weakest on precedent-citing) — trust the shape of each row against the corpus-mix bar above more than any single cell's exact share.
ACT II
One Hand
what one reviewer sets down, alone · plates IV–VI
Interlude · The Three Windows

One question, three windows

The next two plates ask one question through three windows — when a review is read in the order it was written, does position carry information? — and a reader meeting the three analyses in a row may fairly ask what separates them. The answer is the window. Below is the same review three times: each stack is one review's units in written order, coloured by the section of the review form they sit in — strengths, weaknesses, questions, the same colours the plates ahead use for positive, negative, and hedged verdicts. The bright part is what that plate reads; the fine print beneath names the chance level its ×-numbers are measured against. Click a panel to jump to its plate. And one adjustment of expectations, made honestly up front: these three windows mostly find weak or negative results. The strongest order in a review turns out to belong to the review form, not the reviewer; the grammar hunt comes back nearly empty; what survives — who gets the first word of the criticism — is real but modest. The plates keep these results because the absences are informative: knowing that judgment has almost no syntax is worth as much as a syntax would have been.

Plate IV · The Syntax
Corpus: ICLR 2018–2026 · 183,717 initial reviews, 829,684 units

The search for a syntax finds almost none

Plate III sorted every unit into six argument forms. Do the forms care what comes next? For every review, count which form follows which between consecutive units, and compare against a null that preserves the review's own composition — so a common form cannot masquerade as an attractive one.

Fig. 4a — The transition table, as departure from each review's own mix
The grammar hunt comes back nearly empty: no transition beats chance by more than 18% — and the little structure that remains looks like the review template, not a way of thinking.
WHO SPEAKS FIRST, WHO LAST — the flanking OPENS / CLOSES columns ask of the six argument forms what the next plate’s prologue will ask of the twelve subjects: who gets the first word of a review, and who the last.
CELL = LIFT VS. EACH REVIEW’S OWN SHUFFLED MIX · 1.00 = ORDER CARRIES NO INFORMATION
Fig. 4a
How to read
each cell is a lift: how often one argument form actually follows another, divided by how often it would if each review's own units were shuffled — so 1.00 means order carries no information, and if the forms did not care what comes next, the entire table would sit at 1.00. Sienna marks a following that happens more than the review's own mix predicts, verdigris less; the flanking columns give each form's share of opening and closing positions.
Result
the finding is an absence. No cell leaves the 0.80–1.18 band: once a review's own mix of forms is fixed, knowing the current form changes the odds of the next one by at most ±20% — argument order is nearly memoryless. If reviewer reasoning had a real syntax, the way sentences do, cells here would run to several times chance; nothing comes close. What little structure survives: , , and — eight points above chance — while closing shares sit near every form's baseline (largest deviation: norm, +1.7pp).
Interpretation
read even that residue with suspicion: reviews are written into a template — a summary up front, strengths and weaknesses in blocks — and the template alone would produce exactly these traces. The on-balance-first opening matches the summary-section convention, and same-form runs match section blocks. So the honest reading is not "reasoning has a grammar" but the reverse: we went looking for one, found almost nothing, and what remains is plausibly the architecture of the review form rather than of thought. Forms from the Plate III classifier (CV accuracy 0.685); transitions pool 183,717 reviews, 829,684 units.
Plate V · The Itinerary
Corpus: ICLR 2026 · prologue: all 410,586 units in written order · itinerary: 192,085 units anchored to their exact line

Criticism is written in an order

Why this exists
this plate reads the review as a document written in order — in two passes. The prologue takes the widest possible window, the whole review first unit to last; one source of regularity there is known before looking — the ICLR review form itself is ordered, strengths before weaknesses, questions at the end — so much of what the prologue finds is the form’s script, not the reviewer’s, and its captions separate the two where the data allows. The itinerary then strips that scaffolding away: in the 2026 reading every unit records the line of the review it came from, so the [weaknesses] section can be isolated and read on its own — a window inside which the form dictates no order at all. Whatever regularity survives there is the community’s own habit, not the template’s. One honest limit before reading: this is the order of the written document, not necessarily the order of thinking — a review is composed, and a reviewer may reorder, template, or polish before submitting. What the plate can claim is that the writing has strong conventions. (Plate IV asked the same kind of question of the argument forms and found almost none; here the subject matter is what carries the order.) The quantities, in plain terms. Prologue: for each of the twelve objects, an opening lift and a closing lift — how often that object is a review’s first (or last) unit, divided by how often it appears anywhere, so a frequently-raised object cannot look like a favourite opening move just by being common — and the verdict arc: each review with three or more units cut into fifths by position, each fifth’s units sorted by verdict — negative, positive, or hedged, the pooled name used here for conditional, uncertain, and mixed judgments. Itinerary: the bill of faults — a fault is a unit whose cited line falls inside the review’s [weaknesses] section, and for each subject the chart shows where its faults sit within that section, plus how often the subject is written first against random ordering (asked, beside each row, is the share of the subject’s units raised in the [questions] section instead) — and the traveling companions: across consecutive units, which subjects follow one another, and which mostly just share a review.
Prologue · The Script — the whole review, where the form writes first
Fig. 5a — Opening move vs closing move, by object (lift vs base rate)
Reviews open on what the paper is — framing at 1.86× its base rate, novelty 1.57× — and close on clarity (1.81×), related work, and reproducibility.
SAME QUESTION AS FIG. 5a's FLANKING COLUMNS — first and last word, now for subjects rather than argument forms, as a ×chance ratio. the itinerary below re-cuts the opening end inside the weaknesses list alone, where the form imposes no order.
OPENS — how often this subject is a review's first unit, divided by its share of all units; right of the rule = a favourite opening move
CLOSES — the same ratio for the last unit; the vertical rule is ×1, where position tells you nothing
Fig. 5b — The arc of a review
Warm open, hard middle (82% negative at its peak), softened close — an arc the review form itself writes: strengths first, weaknesses in the middle, questions at the end.
Fig. 5a
How to read
one row per object of scrutiny, sorted by opening lift; the horizontal axis is the lift defined in the lede, drawn on a log scale so ×2 sits as far right of the ×1 rule as ×0.5 sits left of it. The legend above the chart names the two marks; hover a row for exact values.
Result
(5.5% of first units against 3.0% of all units), with novelty (1.57×) and method design (1.37×) behind it; the closing end belongs to clarity (1.81×), related work (1.66×) and reproducibility (1.41×). No subject claims both ends, and statistical rigour — a plausible guess for a closing subject — closes exactly at its base rate (1.00×) while avoiding the opening (0.63×).
How much of this is the form?
The whole-review cut mixes the strengths section into the opening, and that matters most for the top row: when problem framing opens a review, 49% of the time it opens as praise — "the paper tackles an important problem" is the classic first strength. Recomputed over negative units only, asking which subject the criticism starts with, problem framing's lift drops to 1.34× while novelty's sharpens to 1.65× (89% of novelty's openings are objections). The closing side barely moves under the same check — clarity 1.81 → 1.77, related work 1.66 → 1.65, reproducibility 1.41 → 1.40 — so closing on clarity is a habit of the criticism itself, the "minor points" tail; and those closing clarity units are asserted, not asked (78% negative, 8% hedged). The opening result is also confirmed below: Fig. 5c makes the same cut independently, never leaving the weaknesses section, and finds novelty leading off the listed faults at 1.62× against the 1.65× here — two cuts, one habit.
Fig. 5b
How to read
the horizontal axis is position within the review in fifths; each curve is the share of that fifth's units carrying one verdict — negative, positive, or hedged (the dashed line; conditional, uncertain, or mixed, as defined in the lede) — pooled over the reviews with three or more units, so the three curves sum to 100% in every fifth. If position carried no information, all three curves would run flat at their overall share; every bend is the finding.
Result
the negative share runs 70% in the opening fifth, peaks at 82% in the middle, and falls to 60% at the close. Praise does something different from a mirror image: 26% of opening-fifth units are positive, collapsing to 6% by the third fifth and 5% in the fourth, then only half-returning (13%) at the close. .
How much of this is the form?
Most of it, and that is the honest reading: ICLR's template asks for strengths before weaknesses and questions at the end. The warm opening is the strengths section passing by; the hard middle is the weaknesses section (93.4% of weaknesses-section units are negative — Fig. 5c documents this); and the close softens less because warmth returns than because assertion gives way to questions — it is the hedged curve, not the positive one, that claims the final fifth. The arc is better read as the form's script, executed in near-unison across the corpus, than as a rhetorical choice each reviewer makes.
Reading it
this describes where in a document a kind of judgment tends to be written down; it is not evidence about the order in which the reviewer thought, or decided.
The Itinerary — inside the weaknesses list, where the form imposes no order
Fig. 5c — The bill of faults: what a reviewer lists first
What a paper is — novelty, framing — gets objected to first (novelty leads the bill at 1.6× chance); how it is measured — statistics, reproducibility — comes last (0.7×).
SAME RATIO, NARROWER WINDOW — the dot column is the same “written first, ×chance” ratio as the prologue's OPENS (Fig. 5a above) — computed here inside the weaknesses list only, so the review form's section order cannot produce it.
Fig. 5d — Traveling companions: which subjects share a review
Mostly company, not sequence: theory fills the reviews it enters, while novelty is said once — .
“Written first, ×chance” — the dot column: one quantity, three states
OPENS — dot right of the ×1 line: written first more often than random ordering (≥1.15×; the lede defines the ratio)
DEFERS — dot left of the line: less often (≤0.85×)
NEUTRAL — near the line (within ±15%)
“Asked” — the bar column: a separate measure, which section
ASKED — bar length: the share of the subject’s units raised in the [questions] section instead of the weaknesses list
“Fault” throughout = a unit whose cited evidence line falls inside the review’s [weaknesses] section — defined by place, not tone (93.4% of them are negative, most of the rest conditional).
NODE — one subject; circle size = its share of the corpus’s units
HALO — dwell: how often a unit is followed by another about the same subject, versus corpus-wide chance; drawn only above ×1.12, thicker = stronger (theory ×1.9 wears the thickest; novelty, below chance, has none)
CHORD — two subjects that follow one another above chance (lift ≥ 1.1 in either direction, ≥ 1,600 successions); thicker = more successions, denser = further above chance
SPARK — animation tracing the chords, faster = stronger pull; hover any node to isolate its connections
Fig. 5c
How to read
each row is one object of scrutiny; the ribbon is where its criticisms sit inside the weaknesses section (left edge = the section's first line, right edge = its last); the ribbon's height at any point is the share of that subject's faults falling at that depth of the section, each row scaled to its own peak — so compare shapes and positions, not absolute heights across rows. The brass tick marks the median position. Everything is computed over the 30,869 reviews with two or more line-anchored faults. One shape is shared by every row and means nothing: the spike at the left edge. A list's first item sits at its start, so each review deposits exactly one fault at position zero — with ~3.9 faults per review, a quarter of all faults (25.0%) are some review's opener. The signal is not the spike but who claims that opening slot beyond their share — the dot column — and the shape after it. If the order of writing carried no information, every ribbon would center on the same spot and every lead-off lift would sit at 1.0; neither happens. Lead-off lift compares how often an object is the first fault listed against its overall share of faults: , with problem framing (1.4×) and method design (1.2×) behind them, while statistics, robustness, and reproducibility are filed toward the end (0.7×). The chip's second number, asked, is a different axis — not where in the weaknesses list, but whether the subject lands there at all — and the two ends line up: novelty, written first, is also the least often softened into a question (9%, the lowest of the twelve), while robustness, written last, is the most often asked rather than asserted (29%). Between the extremes the trend is loose, but the ends are clean: what is written first is asserted; what is deferred is questioned. One candidate reading of that split: questions go where an answer could change something — a missing detail, an absent error bar — while novelty, which no rebuttal can repair, is asserted outright. Read descriptively: the order of the bill mirrors how the community's writing habits arrange criticism — identity-level objections first, measurement-level objections later — not a ranking of importance by any one reviewer. And one rival explanation for the order itself cannot be excluded: reviewers may simply write objections in the order the paper presents its material — novelty and framing live in a paper's opening pages, statistics and reproducibility in its experiments — so the bill may mirror the paper's own table of contents as much as any habit of judgment. One tempting explanation can be checked and set aside: the ICLR 2026 Reviewer Guide contains no novelty-first checklist — its step-by-step opens with clarity, technical correctness, and reproducibility, the very objections reviewers file last — so the observed order is a habit of the community's prose, not an echo of the instructions.
Fig. 5d
How to read
the twelve objects sit on a ring, sized by how much of the corpus they occupy; a halo marks dwell — how often a unit is followed by another about the same subject, against corpus-wide chance (theory 1.9×; novelty 0.72×, the only subject below chance). Chords join subjects that follow one another across 336,206 consecutive-unit pairs more often than independence predicts (lift ≥ 1.1 in either direction, at least 1,600 successions joined): , compute cost with robustness (1.3×), statistics with reproducibility (1.25×). Hover a node to isolate its connections. One apparent paradox is worth disarming: — not because it travels alone but because it travels with everyone. At 21.8% of all units it follows and precedes every subject at almost exactly that subject's base rate (its strongest pairing reaches just 1.03× chance), and a companion to all is a companion to none in particular.
What this order is — and is not
the baseline here is corpus-wide chance, not the stricter within-review shuffle of Fig. 5a, and the difference matters: re-run under Fig. 5a's null, the halos flatten to 0.94–1.29 — subject adjacency is as weak as form adjacency — and even novelty's repulsion inverts (1.20); the small panel beneath the ring draws that collapse, subject by subject. So the ring is mostly a map of company, not sequence: theory reads dense because theory-heavy reviews contain many theory units, and the strongest chords are subjects that share reviews. The fact that survives either reading lives at the review level: 41% of reviews that raise theory raise it again, while 96.5% of reviews that raise novelty say it exactly once — the highest one-mention rate of the twelve. Three candidate readings, offered as hypotheses. Theory's fullness is consistent with proof-checking being sequential work: a derivation read step by step can raise an objection at every step, so one theory reading yields several units where a novelty comparison yields one. At the other pole, the identity subjects — novelty, framing, related work — are said once in over 92% of the reviews that raise them, as if a verdict needs stating, not developing. And the strongest adjacency that survives the stricter null belongs to baselines-and-ablations (1.29): missing-baseline requests read like enumeration, one item after the next. All three readings are hypotheses against the weaker null; under Fig. 5a's stricter shuffle only the baselines adjacency survives, so the ring should be read as a map of company, not of sequence.
Plate VI · The Early Verdict
Corpus: ICLR 2026 · 74,380 rated reviews

The score counts the verdicts — it barely reads the topics

What was done
every 2026 review was cut down to its first k units, and a logistic classifier was trained on that stub alone — the objects those units inspect and whether each came out negative or positive — to guess whether the reviewer's rating was 6 or higher. Sweeping k from one unit to fifteen, and scoring each fit by three-fold cross-validation, shows how much of a review you have to read before your guess stops improving. Because half of all reviews run to five units or fewer, the same sweep is run again on the 2,649 reviews that run to ten units or more, where "does the rest add anything?" is a real question rather than an artefact of reviews running out.
What would be trivial
that the opening beats a coin flip is not a finding — 72% of units are negative and the score tracks negativity, and the target is the same writer's own rating, so panel ① below is the expected part, drawn for scale. The findings are the other three panels: where the climb stops, what it is made of — and what one negative weighs.
Fig. 6a — Verdict readability against units read
The rating is a weighted tally: — the same price order Plate XIV’s tariff will find — yet the count alone does almost all the ranking, and nothing after the ninth unit changes it.
NOT A FOURTH WINDOW — Plates IV–V described where things are written; this plate weighs how much the opening alone already tells you about the score. The first slot's privileged status there becomes a number here.
Fig. 6a
How to read
the four panels share one currency: AUC, the chance the classifier ranks a randomly picked 6-or-above review over a randomly picked lower one — 0.5 is a coin flip, and height or length above the coin-flip mark is the signal. ① is the cumulative climb: the horizontal axis is how many units the classifier was allowed to read, the filled area its lead over chance, the dashed line what the whole review achieves. ② is the same sweep differentiated, on the long reviews only: each bar is the AUC gained by reading one more unit — the first unit's +0.23 is off the scale and stated in the corner; bars turn pale where the gain dies. ③ decomposes the whole review's AUC — verdicts alone, then topics added, then a third bar that abandons the unit features entirely and reads every word (TF-IDF over the full text). ④ asks what one negative unit weighs: a logistic fit on whole-review per-topic negative counts, each bar the pull of one negative of that topic on the score, expressed as a multiple of the average negative's pull — the dashed line is ×1, an average negative.
Result
the findings are ③ and ④, read together. The verdicts buy +0.260 of the +0.279 climb and the topics +0.019 (in the long subset, −0.006 — nothing but noise). Yet the tally is genuinely weighted: one novelty negative counts as 2.4 average negatives — six robustness negatives — and the ordering of these weights reproduces the within-paper tariff of Plate XIV almost exactly (Spearman ρ = 0.90): two designs, one price list. (In the model's own units, a novelty negative is −1.04 log-odds toward a sub-6 rating; the method page carries the details.) Both facts hold at once because so little of the between-review variance rides on composition — weighting the tally beats the plain count by just +0.009 AUC. Topic-blind for ranking, priced for each unit. The words themselves reach 0.822 (TF-IDF, so a floor, not a ceiling, on what the text carries) — more than the coarse profile (+0.043) — and even so, a wide residual stays outside every reading here; Plate XXV will show how much of it is the reviewer, not the review. ② the ninth unit still pays a sliver; from the tenth, nothing — even with fifteen units on the page, the tail is redundant with the head (the few-thousandths jitter is cross-validation noise). ① is the expected part, drawn for scale: ; the pooled flattening at six is partly arithmetic — a median review is over by five — which is exactly why ② is computed on the long reviews alone.
What it shows, and what it does not
the target is the same reviewer's own rating, so this is the internal consistency of a document with its own summary number, not a prediction of anything external — and it is a statement about the text, not about when the reviewer decided: later units largely restate the first, which is exactly what you would write if the conclusion came first. Provenance: ③'s words bar and all of ④ were added as follow-up analyses during review — declared exploratory, committed to publication whichever way they came out. ④'s weights are descriptive regression coefficients on correlated counts, not causal prices; their agreement with the tariff's independent design is what earns them the reading. ⚖ reviews are not interchangeable draws — early-filed ones run harsher and longer, uncontrolled throughout — the tide · overture to act VII
ACT III
The Law
what a criticism actually says — and asks for · plates VII–XIII
Plate VII · The Elements
Corpus: 698,507 negative units across all twelve objects, 2018–2026 · ground → law → remedy, decomposed at the unit level

What a charge actually says

Act II watched one reviewer write — the order, the arc, the early verdict. This act stops watching the writer and cross-examines the writing. Tell a first-year researcher their paper "lacks novelty" — or rigor, or clarity — and they learn nothing. This plate takes such a sentence apart. Because each unit stores the reviewer's observation and their reasoning separately, every negative unit can be split into three: the ground — what the reviewer actually saw in the paper; the rule — the general principle they applied to it, quoted in their own words; and the remedy — what they asked the author to do, if anything. (The courtroom words are used consistently from here on: a negative unit is a denial or a charge, and the principle behind it is a rule or a law.) One more borrowed word: a docket is all the criticism aimed at one of the twelve objects of scrutiny from Plate I, so "the novelty docket" means every criticism of novelty. Pick one below and it opens as ground → rule → remedy. A single criticism can invoke several rules at once, and nothing here forces it into one.

Fig. 7a — The decomposition of a charge: novelty
Pick a docket: every charge decomposes into what was observed, the laws it invokes, and the remedy it demands — .
GROUND → LAW → REMEDY · RIBBON WIDTH = SHARE OF THE DOCKET’S DENIALS · RIBBONS DO NOT CONSERVE
Fig. 7a
How to read
pick a docket above. Left: the grounds — what the reviewer actually observed (named per docket; shares estimated on 40,000-unit samples where a docket exceeds that). Middle: the laws — the definitions the denials invoke, counted only when the reviewer states the rule in words (26–40% of denials do; the rest leave it implicit; a stated denial can invoke several). A law tagged VISITING lives in another docket — the evidence rule visits every court, novelty's own rule sits in six others, and clarity's in three: the jurisprudence of imports is general. Each law card carries its own remedy mix as a thin bar; ribbons on the right carry the same mix outward to the remedy cards, colored by remedy — they represent only the denials whose law was stated, and like the left ribbons they do not conserve. The right column totals everything, and Fig. 7c compares the totals across all twelve dockets. Provenance, stated plainly: cluster structure is data-driven (local embeddings, k-means); merges and names are readings of exemplars — novelty's by the analyst, the eleven extensions drafted by machine readers from the exemplars and reviewed by the analyst; raw clusters ship with the dataset, and the sentence-mining pattern is a prior about normative phrasing, so all stated-law shares are floors. An independent three-reader audit of the borders is recorded in Method §10.
Fig. 7b — The rulebook: novelty's unwritten rules, in reviewers' own words
Six operative definitions of novelty, mined verbatim from 12,789 normative sentences — each clause quotable, dated, and linked to the paper it was written about.
Fig. 7b
Each clause is the sentence closest to its law's center — an actual reviewer's sentence, not a paraphrase written by us (the wording is still the model's tidied version of that sentence; the link leads to the original on OpenReview). Together they are the operative definition of novelty at this venue: distinctness that the author must articulate; more than an increment, more than an assembly, more than a transfer; and, sometimes, explicitly indexed to the venue's rank. One clause is deliberately kept though it is not a law of novelty at all: — a jurisdiction switch worth seeing. The rulebook shown is novelty's — the pilot docket, read deepest; the other eleven carry their laws in Fig. 7a, and Plate VIII opens one clause — the combination rule — all the way down.
Reading it
an audit stands behind the borders: three independent machine readers, given raw denials and no taxonomy, re-derived these families — and surfaced two finer clauses the clustering had folded away (the false-priority accusation; the contribution that cannot be located at all), recorded in Method §10.
Fig. 7c — What each charge asks of you
Every docket has a remedy signature: clarity wants writing (90%), baselines want experiments (80%), reproducibility wants disclosure (76%) — and novelty, alone, wants nothing at all 39% of the time.
BAR = SHARE OF THIS OBJECT’S DENIALS, BY REMEDY REQUESTED
Fig. 7c
How to read
for every object of scrutiny, its denials split by what the reviewer asked for: articulate (writing and positioning), substantiate (new experiments and analyses), redesign (change the method itself), disclose (details, code, data), or nothing. The signatures are almost caricatures of each docket — . The outlier is the plate's oldest finding made general: — . One artifact was checked and ruled out: novelty complaints are not terser than others (slightly longer, in fact), and the gap persists inside every quartile of unit length (novelty offers a fix for 60–61% of its denials in each quartile, against 76–94% for the other objects). “Redesign” , which is itself worth knowing: reviewers ask authors to argue, prove, and disclose far more often than to change the method.
Reading it
remedy shares for the eleven extended objects are estimated on 40,000-fix samples per object where fixes exceed that, and each docket's fixes are read through eight clusters — so an exact 0% means rarer than the clustering resolves, a floor rather than a measured absence.
Fig. 7d — The interchangeable syllogism
If a charge were a syllogism, this observation would summon that rule, and that rule would fix the ask. Inside the novelty docket the current runs — but barely: . The stages of the argument are nearly interchangeable parts.
NOT A CONTRADICTION OF FIG. 4a — THAT WAS THE ORDER OF WRITING; THIS IS THE LOGIC INSIDE ONE CHARGE. AND ONE NUMBER RETURNS BY ANOTHER ROAD: 39.4% OF THESE 30,447 DENIALS ASK NOTHING — FIG. 7c FOUND 39% ON A DIFFERENT PIPELINE.
RIBBON WIDTH = SHARE OF THE 30,447 DENIALS TAKING THAT PATH · COLOR = ×CHANCE (WARM ABOVE, COOL BELOW, PALE ≈ 1.00 = INTERCHANGEABLE) · CRAMÉR’S V 0.14 & 0.14
Fig. 7d
How to read
every negative novelty unit in the nine-year corpus, cut into its three stages — what was seen (ground), the rule invoked (warrant), what was asked (demand) — each stage clustered on its own, then the flow between them counted. Each ribbon is real traffic — its width, the share of denials taking that path; its colour, that path’s traffic as a multiple of chance (warm above, cool below; hover any ribbon for the exact figures; the dashed bar is the empty ask). If the syllogism were a real circuit, a few ribbons would run hot and the rest would vanish — Plate XIII’s coupled rule-pairs reach 3.6×. Instead nearly every ribbon is pale: each observation fans out over all the rules at close to their base rates, and each rule fans out over the asks the same way. In the corpus’s own words, the stages sound like this — a ground: “Method resembles a combination of existing approaches”; a warrant: “Combining existing techniques without a conceptual leap is not a strong contribution”; and after that warrant, the modal ask is nothing at all.
Result
the coupling is real — χ² rules out independence at this n — and everywhere faint: Cramér’s V ≈ 0.12–0.14 at both stages, both grains. . The channels that do carry current tell one story: , and — while rigor-doubt reasoning almost always asks for something (25% empty). The novelty syllogism is less a chain of reasoning than a bin of interchangeable parts: any observation licenses nearly any rule, and the rule barely narrows the ask.
What was done
the observation, reasoning and suggested-fix fields of all 30,447 negative novelty units (2018–2026, forum track) embedded separately (bge-small) and k-means clustered — 10 grounds, 10 warrants, 8 demands, seed 46 — then merged to the five/four/three named stages shown, by reading exemplars; raw clusters and matrices ship in the dataset. The χ² and V above are quoted at the raw grain, before any merge, against the product of marginals — the marginal-preserving shuffle null. The ask stage counts “nothing asked” as an outcome, so its V includes the decision to demand at all.
Reading it
faintness is the finding — this figure must not be quoted as “criticism has a logic-grammar after all.” A 1.9× channel over a 1.00 floor is a lean, not a law. And the stages are the model’s distillation of the reviewer’s words, read one unit at a time: a distiller that phrases rules generically could smooth real structure away, so these lifts are floors, not ceilings — what survives distillation is faint; what the reviewer privately walked through is beyond this instrument.
Plate VIII · The Combination Clause
Corpus: ICLR 2018–2026 · 4,316 combination-rule sentences from 4,262 units

What “mere combination” actually says

Plate VII found the combination rule — assembling known components is not novel unless the integration itself is. This plate opens the clause: what, concretely, is the unless? And what is being called a combination in the first place? Sub-clustering the rule’s own sentences answers the first; the units’ observations answer the second.

PLATE VII’S SPLIT, RUN INSIDE ONE LAW — Plate VII cut every charge into ground → law → remedy. This plate takes a single law and cuts it the same way: Fig. 8a reads the law side (the rule’s own sentences, hence % of statements), Fig. 8b the ground side (what the same charges observed, hence % of charges).
GROUND → ● LAW → REMEDY — THE CELL THIS FIGURE OPENS
Fig. 8a — The redemptions: what would make a combination count
Six distinct ways out, — added value, new insight, differentiation, a new mechanism, an isolated contribution, innovation beyond the integration — and two clauses that offer none.
● GROUND → LAW → REMEDY — THE CELL THIS FIGURE OPENS
Fig. 8b — The referents: what is being called a combination
In 46% of charges the assembled parts are never named: the reviewer says the work combines existing methods without saying which.
BAR = SHARE OF COMBINATION CHARGES POINTING AT THIS KIND OF REFERENT
Fig. 8
How to read
left, the rule drawn as its own grammar — a railroad of the 4,316 combination-rule sentences, sub-clustered and named. Every statement enters at the same spine — a combination of known parts is not novel — and each track is one family of the rule, its weight the family’s share. The six brass tracks are genuine exception clauses — each names the thing whose presence redeems the assembly — and they rejoin the main line: novel after all. ; 24% of statements offer no exit. Hovering a track lights its card below, where each clause carries its definition and a reviewer’s own sentence. An explicit conditional word (unless, without, only if) appears in 43.6% of the statements — a floor, since the track families are read from a sentence’s content rather than its grammar, and together the six exception tracks hold 76%. For a rejected author the left column is the actionable reading of “just a combination”: . Right, the same units’ observations grouped by what they point at: , , , . The 46% is measured on the distilled observations, so it was audited against the source reviews: in a 30-case sample, wherever the source charge could be located, its specificity matched the observation’s — generic charges were generic in the reviewer’s own words too, with one ambiguous exception — so the share is not an artifact of summarization. Clusters are k-means over sentence embeddings, names drafted by a machine reader over exemplars and reviewed by the analyst (method § 10); raw clusters ship with the dataset.
Plate IX · The Unnamed Precedent
Corpus: ICLR 2018–2026 · 30,824 negative novelty units

The judgment that needs no referent

Plate VIII opened one clause of the novelty rule from the inside — what redeems “mere combination.” This plate opens a different clause the same way: every novelty objection implies a comparison — this work against the prior it allegedly repeats. We went looking for the canon: the named prior works that kill novelty. The search came back nearly empty, and the emptiness describes the standard itself. A judgment that cites nothing is not automatically careless — novelty can be judged from the reviewer’s internal map of the field, which is much of what expertise is — but a judgment made that way cannot be checked from the text, and its burden of proof moves to the author. What follows measures that mode of judgment: how common it is, and which way it has moved as the venue grew.

Fig. 9a — Of every 100 novelty objections: who names the prior?
12 name a specific work. 25 gesture at "prior work." 63 assert un-novelty with no referent at all — and , under both instruments tried.
EACH MARK = 1% OF THE 24,233 LOCATABLE NOVELTY OBJECTIONS
Fig. 9a
How to read
each mark is 1% of the 24,233 negative novelty units whose charge could be located in its source review (79% of 30,824). The unit only locates the charge; the measurement runs on the reviewer’s original words — a citation regex over every sentence mentioning novelty, plus two sentences either side. , , . The line below tracks the named share by year: — as the venue grew, the charge became more anonymous, not less. One pair, side by side — both distilled units quoted verbatim from the shipped clusters: a lit mark reads “Prior works (Cohen et al. 2023, Choi & Chung 2019) address similar concepts; the paper fails to clearly distinguish its contribution” (2025) — the citation is extractable; a dark mark reads “The method appears to be an incremental combination of existing methods. The contribution lacks fundamental novelty compared to prior art” (2026) — no referent at all, and the burden of naming the collision passes to the accused.
Two instruments, one reading
an earlier version ran the same regex on the distilled units instead and found 5 / 29 / 66. The naming level is instrument-sensitive — the tight unit reading is a floor, and the ±2-sentence window leans high because it can absorb a neighbouring criticism’s citation — but the two facts this plate rests on hold under both: roughly two thirds of novelty objections carry no referent at all (66 tight, 63 generous), and the naming rate halves from 2018 to 2026 under either reading (8.4%→3.5% tight; 20%→11% generous). Nor is there a canon among the named: under a deliberately generous key (surname + year, which merges distinct papers), the most-cited key reaches 21 of the 23,251 papers these objections were filed against. The accusation of un-novelty is structurally unnamed: the burden of locating the colliding prior work is left with the accused. The regex misses prose references ("the BERT paper"), so all lit shares are undercounts.
Plate X · The Formulary
Corpus: ICLR 2026 · 293,671 negative units, embedded locally · nearest twin always from a different paper

Criticism has a common coin

Take any criticism and search every other paper’s reviews for the most similar one — its nearest twin. Some criticisms have no close twin: they could only have been written about this paper. Others recur almost word for word across many papers — the field’s standard formulas. Finding formulas is not an accusation: a shared formula is how a discipline applies the same standard to everyone ("release the code" has to read the same everywhere to be a standard at all). The question is which criticisms are formulas and which are written for the paper at hand — and Fig. 10b’s answer is that the two kinds carry very different weight. After three plates of nine-year corpora, the window narrows again: the embedding search reads ICLR 2026 alone — 293,671 negative units.

Fig. 10a — The spectrum: how close is your criticism's nearest twin?
Most criticisms have a familiar twin in someone else's review — .
← LESS SIMILAR · NEAREST-TWIN COSINE SIMILARITY · MORE SIMILAR →
Fig. 10b — Formula share by object, against what the objection costs in score
The objections that cost the most score are the hand-written ones (novelty: 1.3% formula share); the cheap ones are standard formulas (reproducibility: 8.7%).
NOT THE OPPOSITE OF PLATE IX — unnamed is not formulaic. Novelty’s charges almost never cite a precedent (Plate IX), yet their wording is the least die-struck here: a charge can be tailored to this paper and still point at nothing.
Fig. 10a
How to read
every negative unit's observation is embedded and matched against every unit from other papers; the curve is the distribution of that nearest-twin similarity. — familiar, but not identical; , and . And the formula share rises with the score — 3.6% at ratings 0–2 but 6.8% at a 10 — because an admiring reviewer must still fill the weaknesses box, and fills it from the formulary.
Reading it
one honest caveat carries this whole plate: what is compared is the model's tidied restatement of each criticism, not the reviewer's raw sentence, and that tidying pulls different phrasings closer together — so read the absolute levels as upper bounds, and trust the comparisons (like the rise with score), which all pass through the same instrument.
Fig. 10b
How to read
one line per object of scrutiny, strung between two poles. The left pole is how formulaic its criticisms are — the share whose nearest twin clears 0.90. The right pole is what raising the objection costs in score — how much lower a reviewer rates a paper when they raise it, compared with the other reviewers of the same paper — the tariff, a quantity Plate XIV will define and defend four plates ahead; its values are borrowed early so the two can be read together. If formulaic and expensive went together, the lines would run level; instead nearly every line crosses, and the crossing is the finding. The inversion holds at every cutoff tried (0.85, 0.90 and 0.95 all put the same four objects at the bespoke end; clarity tops the formula end at 0.85, reproducibility at the other two — and at 0.95 related work edges past clarity; 0.90 shown): the objections that cost points are — while the free objections are . The expensive criticism is written for this paper; the cheap criticism could have been written about almost any paper.
Reading it
the tariff pole is borrowed from Plate XIV's within-paper 2026 regression and carries that plate's caveats (bootstrap intervals, single-year scope); a coefficient whose interval crosses zero is greyed on the figure rather than colored as cost or gain.
Fig. 10c — Twin exhibits: the same criticism, two different papers
Three tiers of twinhood, descending Fig. 10a’s guides — brass marks the words that differ; in the top tier, almost nothing.
BRASS RUNS = THE STRETCHES WHERE THE TWO PAPERS’ REVIEWS DIFFER
Fig. 10c
How to read
each card holds one criticism as read from two unrelated papers' reviews, facing each other across the pair's similarity; brass marks the words that differ between the two, so the amount of brass on a card is the distance between the twins. The cards descend Fig. 10a's guides: ; ; . Both quotations are the model's tidied restatement, and each links to its source review on OpenReview so the paraphrase can be held against the original.
Reading it
a near-identical pair is not evidence of a careless reviewer. Two papers can share a real defect, and a discipline needs portable standards — what the pairs show is which criticisms are portable, and that the portable ones are concentrated in the cheap dockets.
Plate XI · The Repair Manual
Corpus: ICLR 2026 · 307,388 units carrying a suggested fix · clustered locally, named by hand

What reviewers actually ask for

When a review is split into units, each criticism keeps whatever concrete suggestion the reviewer wrote next to it — the same fix field Plate VII read as remedies. 85% of criticism — and three quarters of everything reviewers set down — arrives with such a repair attached. This plate asks what those repairs actually say: 80,000 of them are embedded and clustered into a working taxonomy of the discipline's demands. Read at the family grain of Plate VII, one kind dominates — over two thirds of everything requested is work of research: new evidence in some form.

Fig. 11 — The sixteen repairs, by share of all suggested fixes
— and its whole family, the work of research, is over two thirds of everything asked.
THE SAME FIELD AS PLATE VII’S REMEDIES — Fig. 7c read the fix field at four-remedy grain, per docket, over nine years; here the same field is clustered into sixteen named types across 2026 alone, and the colors carry Plate VII’s families. Once again, no type asks for a redesign.
BAR = SHARE OF THE 80,000 SAMPLED SUGGESTED FIXES
Fig. 11
How to read
the thin bar on top is the headline: the three-family composition of everything asked. Below it, each bar is one repair type's share of the sampled fixes, grouped and colored by Plate VII's remedy families; click or tap a row to open a verbatim exemplar and the objects the repair most often attaches to. The family assignment is the analyst's, with two judgment calls recorded: error bars count as work of research (they must first be computed), surface fixes as work of writing.
Result
— but a single cluster's share depends on how finely the taxonomy is cut (), so the stabler finding is . And one absence, drawn as the empty dashed shelf at the bottom: , the same absence Fig. 7c found at coarse grain. On first view the figure assembles as sediment: each falling grain is one sampled ask, settling into its shelf — and nothing ever lands on the redesign shelf.
What was done
of 2026's 410,586 units, 307,388 carry a suggested fix; 80,000 of those were sampled, embedded (bge-small) and clustered (k-means, k=22, seed 7), then merged to sixteen types and named by hand from each cluster's most distinctive words (class-based TF-IDF) and most central examples.
Reading it
that absence is a fact about what reviewers write, not about what they think. A repair is an ask the authors could plausibly act on within a rebuttal, and "start over" is not — which may be exactly why the harshest criticism arrives with no repair at all (novelty's 39% in Fig. 7c). The taxonomy's sixteen-to-three mapping and the two recorded judgment calls are this figure's error bars; any single cluster's share moves with how finely the 22 raw clusters are cut, though the family totals do not.
Plate XII · The Verdicts
Corpus: ICLR 2026 · 410,586 units

Where scrutiny turns hostile — and which criticism arrives with a way out

Fig. 12a — Verdict mix per object, sorted by negativity
Positioning work — — is judged almost only in the negative.
NOT NEW DATA — the marginal of Plate I's diagram: the same units' object → verdict edge, re-sorted as bars.
EACH ROW = ONE OBJECT’S UNITS, HUNG ON A SHARED AXIS BY POLARITY
Fig. 12b — Criticism that arrives with no fix at all
37% of novelty criticism offers no way forward — two to four times any other object.
FIG. 7c’S "NO REMEDY", IN THE 2026 WINDOW — the same fix field Fig. 7c read as remedies over nine years (novelty: 37% without a fix here, 39% there — Fig. 10 read the objections themselves, not their fixes); drawn in the same color as 7c’s no-remedy band.
% = OF THIS OBJECT’S CRITICISM, HOW MUCH ARRIVES WITH NO SUGGESTED FIX
Fig. 12
How to read
in 12a every object hangs on one shared axis: the sienna wing reaching left is the share of its units that criticise, the green wing reaching right the share that praise, and the dim band between them everything in between (uncertain, mixed, conditional — hover for the split; the full five-way mix is Plate I's). The further a row reaches left, the more that topic exists only to be criticised. In 12b each track is 100% of the object's criticism; the dark band is the share arriving with no suggested fix, and the dim remainder arrives with one — so novelty's band is the largest, yet still only a third of its criticism.
Result
, and . A second quantity was checked and retired from the chart: the share of units grounded in the reviewer's own explicit words is flat across all twelve objects (17–22%), so it is reported here rather than drawn.
What was done
every 2026 unit is grouped by the object it inspects. 12a splits each object's units by the verdict they reach; 12b asks two further yes/no questions of each object's negative units — did the reviewer write a concrete suggested fix alongside the criticism, and does the unit quote or point at specific text in the paper rather than resting on the model's inference.
Reading it
"comes with a fix" is a property of what the reviewer wrote, not of whether the problem is fixable, though for novelty the two are hard to separate — a reviewer who thinks a contribution too small often has no repair to offer.
Plate XIII · The Charge Sheet
Corpus: ICLR 2018–2026 · 120,540 of 191,946 reviews state a rule in words

Which laws are filed together

Plate VII collected the rules reviewers state in their own words — proofs must be rigorous, notation must be defined, novelty requires differentiation from prior art. A single review usually states several. This plate asks: which rules tend to appear in the same review together — and what that pattern says about the person writing. The logic is simple: if each rule appeared only because of the paper’s own faults, any two rules would share a review at plain chance rate, so every departure from chance points at something else. Departures are measured with the ×chance lift used throughout the atlas — 2× means a reviewer who states one rule states the other twice as often as chance — across all 191,946 initial reviews. Two kinds of departure appear. Rules about rigour and presentation cluster in the same review, for two reasons at once: a paper that earns one such complaint tends to earn the others, and a careful reader is careful about everything (the two are told apart by checking the other reviewer of the same paper — the caption walks through it). The novelty complaint does the opposite: a reviewer who has written “not novel” once does not write it again under another heading, while the other reviewer of the same paper is just as likely to write their own. The clustering is partly the paper’s doing; the once-only habit belongs to the reviewer alone. In the figures below the rules are called laws and a review’s set of them its charge sheet — the atlas’s courtroom image, and nothing more.

Fig. 13a — The full grid: every pair of rules, filed together or apart
Most pairs run warm (82% exceed chance; median ×1.19) — a longer, more thorough review files more of everything — which is what makes the violet cells the exception: novelty’s clauses push each other off the sheet, down to ×0.66, against the tide.
SAME ×CHANCE IDEA AS FIG. 5d's CHORDS — but the pair here is two rules co-present anywhere on one review's charge sheet, not two subjects succeeding each other in its text.
Fig. 13b — The extremes, read closely
— each with its 95% interval and its cross-reviewer twin.
FILED TOGETHER — THE TOP TEN
NEVER ON THE SAME SHEET — THE BOTTOM EIGHT
Fig. 13
How to read
in 13a, rows and columns are the 61 stated rules, grouped by docket; a brass cell means the two rules land on the same review more often than chance, a violet cell less often, and the barely-tinted cells lack the 60 joint filings needed to be read (same-docket pairs are excluded throughout — two rules about the same object co-occur trivially). Two textures carry the finding. The base tone is warm — the median cross-docket pair sits at ×1.19 and 82% of readable cells exceed ×1, because a longer, more thorough review files more of everything — so the deepest brass (the clarity–theory–statistics neighbourhood, up to ×3.6) and every violet cell are departures read against that tide. The violet cells concentrate where one novelty-flavored clause meets another: filing one pushes the other off the sheet, in the teeth of the verbosity that lifts every other pair. 13b names the extremes — on the matrix they are the white-ringed cells; 3.6× means a reviewer who states one of the pair states the other 3.6 times as often as independent filing would produce (universe: all 191,946 initial official reviews). The faint band behind each pair of dots is a 95% interval from resampling whole papers — every attracted pair’s interval stays above ×1.25 and every repelled pair’s below ×0.91, so the extremes are not sampling luck. Because each docket’s sentences were sampled independently, absolute co-filing rates are thinned, but the ratio between them stays accurate, so only ratios are shown. The attracted pairs are almost all form-and-rigor pairs — , . The repelled pairs are almost all novelty-flavored: .
Paper or reviewer?
a within-review lift cannot by itself say whether the pull follows a careful reviewer or a paper that invites both complaints, so every displayed pair was re-measured across two different reviewers of the same paper (the grey column; test declared before results). The attractions are a mixture: the median attracted pair keeps ×1.44 across reviewers — the paper’s contribution — against ×1.87 within one review, and the excess is the reviewer. (The two reviewers of one paper also share a year; re-basing chance within each year moves the cross-reviewer medians by 0.01, so the shared year contributes essentially nothing.) The repulsions are not a mixture: across reviewers the median is ×0.97, no suppression at all, against ×0.75 within one review — filing one novelty clause suppresses the next only inside the same review. The one-stamp habit is the reviewer’s, not the paper’s. (Cross-docket pairs cannot share a unit, so double-labelled sentences contribute nothing to either list.) One residual the control cannot remove: units are the model’s distillation, read one review at a time, so a machine reader’s internal consistency could contribute to the within-review excess on the attraction side. On the repulsion side that bias runs the other way — an extractor’s consistency can only create positive correlation — so the one-per-review suppression is, if anything, understated.
ACT IV
The Tariff
what a criticism changes · plates XIV–XV
Plate XIV · The Price
Corpus: ICLR 2026 · 74,373 rated reviews of 19,467 papers  ·  standards & stated rules: the nine-year corpus — 384 thousand reviewer pairs, 120,540 rule-stating reviews

What a fault costs, and whose law decides

A review ends in one number. This plate asks three questions that number hides. First, what does each objection cost — how much lower does a reviewer score a paper when they raise it? Comparing reviews of different papers would answer nothing: weak papers attract both criticism and low scores. So the cost is measured between reviewers of the same paper, also holding fixed how much criticism each review contains overall — what survives is the price of that particular objection, the plate’s tariff. Second, is a fault’s price tied to whether it can be repaired — do the costly objections arrive with a prescribed fix, or without one? Third, when two reviewers examine the same object of the same paper, does their agreement depend on the reasoning standard each brought to it — the laws of Plate VII? And fourth, one level finer and nine years wide: when a criticism states its rule in words, does the stated rule set the price — or does the same rule cost differently depending on where it is aimed?

Fig. 14a — The tariff: the price of each objection, paper held fixed
Holding the paper fixed, a novelty objection travels with well over twice the score deficit of any other charge.
THE RULER ITSELF STANDS TRIAL LATER — every price in this act is measured in rating points; how much of a rating is signal and how much is noise is Act VII’s question (Plates XXIV–XXV). The within-paper comparison used throughout here is built to need no cross-paper comparability from that ruler.
AXIS: HOW MANY RATING POINTS LOWER (−) OR HIGHER (+) A REVIEW SCORES THAN ITS CO-REVIEWS OF THE SAME PAPER, WHEN IT RAISES THIS OBJECTION · THE RATING SCALE RUNS 2–10
Fig. 14b — Price against repairability
The expensive charge is hand-made against this paper; the cheap ones are prescriptions.
ONE AXIS BORROWED, ONE NEW — the horizontal axis is Fig. 13b’s measurement, complemented: 13b drew the share of each object’s criticism arriving with no fix (novelty 37%, the rest 8–18%); here the same objects sit at their with-fix share. What is new is the vertical: Fig. 14a’s tariff. This figure adds no measurement — it joins two already made.
Fig. 14a
How to read
each solid point is the regression coefficient for "this review raised at least one negative unit on the object," with each paper's own mean rating subtracted out so only within-paper differences remain, controlling for how many negative units and how many units the review contains overall; whiskers are a 95% interval from resampling whole papers. The hollow point is the naive gap between reviews with and without the objection, papers not held fixed — the distance between the two points is the part explained by which papers attract the objection and by the extra criticism that travels with it. Holding the paper and the criticism volume fixed, on the 2026 rating scale (which runs 2, 4, 6, 8, 10 — even numbers only) — — while . Robustness and compute-cost objections carry a positive sign: they are the objections of reviewers who found nothing worse to say. Read these as associations, not effects. Nobody assigned the objections at random: a reviewer who is already unconvinced is the one who raises novelty, so the number measures which objection travels with a low score, not what raising it would do. ⚖ reviews are not interchangeable draws — early-filed ones run harsher and longer, uncontrolled throughout — the tide · overture to act VII

Fig. 14b — Each object placed by the share of its negative units that arrive with a prescribed repair (horizontal) against its within-paper tariff (vertical). Circle area tracks how many criticisms the object drew, fill repeats Fig. 14a's color — red = costs rating points, green = travels with higher ones, gray = interval crosses zero — and each label is inked in its circle's color, with a thin line tying the two — novelty alone in white, the plate's finding. The bill collects in one corner: the one objection that rarely comes with a fix — — is also by far the costliest. Cheap objections are prescriptions; the expensive one is a sentence. The unlabeled points crowding the near-zero center — — are the objections that neither cost nor spare; hover any point for its numbers.

Fig. 14c — Same bench, different law
Same paper, same object — and still, one split in three among pairs holding different standards travels with the standards alone.
A PAIR = TWO REVIEWERS OF THE SAME PAPER WHO BOTH EXAMINED THIS OBJECT — THE PAPER CANNOT EXPLAIN WHAT FOLLOWS. EACH CRACK SPREADS FROM THE CENTER LINE: LEFTWARD, HOW OFTEN SAME-STANDARD PAIRS STILL REACH OPPOSITE STANCES · RIGHTWARD, THE ADDED SPLIT WHEN THEIR STANDARDS DIFFER
Fig. 14c
Where to look
the top row is the finding, and the bracket marks it. Its brass wing: . Its sienna wing: pairs who argued from different standards, 30.1% — half again as often, and the bracketed +9.7 points is the split added by the standards alone: of the splits among different-standard pairs, one in three is excess over the shared-standard rate (in agreement terms, 79.5% against 69.9%). The paper cannot be the explanation — both reviewers read the same paper. The rows beneath repeat the reading object by object: every sienna wing is positive.
How to read
every crack spreads from the shared center line — brass leftward for the same-standard split rate, sienna rightward for the addition when standards differ — so both wings compare directly across rows, and rows are sorted by the sienna wing.
What was done
across all nine years, take every pair of reviewers who examined the same object of the same paper in their initial reviews; a reviewer's stance on the object is whether the majority of their units on it were negative, and their standard is the reasoning law they used most.
Reading it
the — read their exact sizes loosely. Splits are rare in absolute terms because criticism dominates the corpus — most pairs condemn together — which is why the sienna addition, not the whole crack, is the quantity to read. And stance and standard are extracted from the same few units, so consistency of a reviewer's writing can contribute to the gap. The choice of standard is not prior to the judgment — it is part of it: which law a reviewer reaches for already carries most of the verdict they are about to give.
THE FIVE FIGURES, READ TOGETHER — hold the paper fixed and most objections cost nothing (14a); the one objection that does cost is the one that rarely arrives with a repair (14b); two reviewers of the same object disagree half again as often when they argue from different standards (14c); and the stated rule prices by where it is aimed, not by what it says (14d, 14e). Together they say one thing: what a review writes as its reasons looks less like the input that produced the verdict than like a part of the verdict itself. The repairable criticism is homework that leaves the score alone; the unrepairable charge, and even the choice of standard, travel with the judgment. Association throughout, never causation — but the same picture, three ways.
The Law’s Price — the same rule, priced docket by docket over nine years
Why this exists
a natural model of refereeing goes: the reviewer finds a rule violated — "claims must match their validation", "the contribution is too incremental" — and the penalty follows from the rule that was broken. If that model were right, the same rule should cost about the same wherever it is applied — a prediction that can be read straight off the chart below. (Stated plainly for provenance: the measurement is descriptive and was not designed as a test of this model; the model is offered as the reader’s yardstick, added at review.)
What was done
a minority of criticisms state their reasoning rule in words — rare enough that all nine years, 2018–2026, are pooled, where the figures above read 2026 alone. For each stated rule, in each docket where it appears (a docket is one object’s file: the clarity docket holds criticism of the writing, the novelty docket of originality), the measure is how much lower the rule’s invoker rated the paper than the same paper’s other reviewers — in standard deviations (σ) of that year’s rating scale, because ICLR changed scales four times over these years. The test needs two things held fixed: the paper — handled by comparing each invoker to the same paper’s other reviewers — and the rule itself, which is why the test case is the evidence rule: it is the only rule that appears in all twelve dockets, the only one that lets the docket change while the rule stays put.
What it will show
the prediction does not hold. With the volume of criticism held fixed, the same rule keeps a −0.30σ deficit in the framing docket and none at all in clarity (14d), and the expensive docket’s bill is run up by its own rules, not by imported ones (14e). The reason written down does not set the score — what was doubted does. Association, not causation, as everywhere in this act — but it is the tariff’s lesson (Fig. 14a), one level finer: the stated reason reads less like the cause of the judgment than like an explanation attached to it.
Fig. 14d — One law, twelve courtrooms: the evidence rule’s tariff by docket
Aimed at what the paper claims to be (framing), the same rule keeps a −0.30σ deficit even with the review’s criticism volume held fixed; aimed at how it is written (clarity), it costs nothing.
THE RULE IS IDENTICAL IN EVERY ROW — ONLY THE DOCKET CHANGES. AXIS: THE INVOKING REVIEWER’S RATING MINUS CO-REVIEWERS’ MEAN FOR THE SAME PAPER, IN σ — ONE STANDARD DEVIATION OF THAT YEAR’S RATING SCALE · 0 = NO GAP
Fig. 14e — Inside the novelty docket: the price of each of its rules
Novelty’s own rules — about combining or extending existing work — cost the most (−0.44 to −0.46σ, volume-adjusted); the borrowed evidence rule costs well under half that (−0.19).
SAME AXIS, NOVELTY DOCKET ONLY — BRASS = NOVELTY’S OWN CLAUSES · RED = THE BORROWED EVIDENCE RULE, PRICED DOCKET-BY-DOCKET AT LEFT
Fig. 14d & Fig. 14e
Where to look
the two ends of the left panel: — both marked on the chart; the framing−clarity gap is −0.29σ (95% CI −0.38 to −0.20, resampling whole papers). In the right panel, — the ×2.4 marked on the chart.
How to read
each filled point is the mean within-paper score deficit of reviews invoking that rule — how much lower the invoker rated the paper than its other reviewers did — with the review’s criticism volume regressed out (each negative unit −0.37σ, each unit +0.27σ), so "reviews invoking this rule here are simply harsher overall" cannot explain what remains; hollow circles are the gap before that adjustment. The adjustment was declared during review as a robustness check, committed to publication either way: it shrinks every deficit but preserves the order. Whiskers ±1 standard error; everything in standard deviations of the year’s own rating scale so four scoring regimes can share one axis. The dashed line is what the naive model predicts: if the rule alone set the score, every dot would sit on it — the spread of the dots is the finding.
Result
the price order tracks what the docket concerns: the evidence rule is expensive where the paper’s identity and credibility are at issue — framing, reproducibility, related work, novelty — and cheapest in clarity, where only the writing is. The rule is the same everywhere; what it costs depends on where it is applied. And inside the novelty docket, the expensive rules are novelty’s own — the increment and combination rules, more than twice the price of the borrowed one. Same association-not-causation reading as the tariff plate: the rule a reviewer reaches for is part of the verdict, not prior to it.
Plate XV · The Consequence
Corpus: ICLR 2026 · decisions joined at display time

Which criticism travels with rejection

Why this exists
the decisions were deliberately kept out of the reading pipeline, so the units and the outcomes are independent records. Joining them is the first look at decisions — accept or reject, rather than the scores the tariff plates used — and at whether the criticism has anything to do with them.
What was done
of the 19,474 reviewed ICLR 2026 papers, the 13,704 that carry a decision and were not withdrawn are split twelve times over — once per object of scrutiny — into those that drew at least one negative unit on that object and those that drew none, and the acceptance rate of each side is compared. Nothing is held fixed, so this is a comparison between different papers, not within one. The closing figure counts the same join sideways — not which object drew criticism but on how many.
Fig. 15a — Acceptance rate with (●) vs without (○) negative units, per object
Novelty criticism travels with rejection — a ; compute-cost and robustness criticism travel with nothing.
SAME QUESTION AS THE TARIFF, A DIFFERENT COIN — the tariff (Plate XIV) priced an objection in rating points with the paper held fixed; here the outcome is the paper's decision, with nothing held fixed. And unlike Fig. 5a's dumbbells, the vertical rule is not ×1 chance — it is the venue-wide acceptance rate. One more thing these venue-wide gaps hide: Plate XXVIII will split this same indicator by research territory, where the signatures differ far more than they do here.
EACH ROW SPLITS THE 13,704 DECIDED PAPERS BY ONE OBJECT: ● = ACCEPTANCE RATE OF PAPERS DRAWING ≥1 NEGATIVE UNIT ON IT · ○ = OF PAPERS DRAWING NONE. IF DECISIONS HAD NOTHING TO DO WITH THE CRITICISM, ● AND ○ WOULD COINCIDE ON EVERY ROW
Fig. 15a
Where to look
the top row and the bottom two: .
How to read
one row per object; the filled dot is the acceptance rate of the papers that drew at least one negative unit on that object, the hollow dot the acceptance rate of those that drew none, and the bar between them is the gap; the thin bar under each filled dot is its 95% interval.
Result
novelty shows the largest gap, −10.6 points (95% CI −12.1 to −9.0, resampling whole papers), with ; compute cost (CI −1.7 to +1.4) and robustness (−0.9 to +2.4) show no gap at all — papers criticised on those grounds are accepted at the same rate as papers that are not.
Reading it
association only, and the direction is not established: reviewers raise novelty on papers they are already unconvinced by, so this measures which criticism keeps company with rejection, not which criticism causes it. One thing it is not, though: an artefact of the pipeline. The decisions were never shown to the models that read the reviews, and are joined to the units only here. ⚖ reviews are not interchangeable draws — early-filed ones run harsher and longer, uncontrolled throughout — the tide · overture to act VII
THE SAME JOIN, COUNTED SIDEWAYS — BREADTH OF CRITICISM VS ACCEPTANCE — IS EXPECTED BY CONSTRUCTION, SO ITS FIGURE HANGS IN THE NULL CABINET ↓
Fig. 15b — The legible verdict: how much of the decision is written in the criticism
The atlas’s front-page question, asked literally: can accept-or-reject be read from the criticism alone? — and the tally’s share has been falling for nine years.
THE SAME QUESTION AS PLATE VI, ASKED OF THE VERDICT — PLATE VI READ ONE REVIEW’S SCORE FROM ITS OWN TALLIES (AUC 0.78 TALLY / 0.82 TEXT, 2026). THIS FIGURE READS THE PAPER’S DECISION FROM THE WHOLE PANEL’S PRE-REBUTTAL CRITICISM, NINE YEARS DEEP. THE DECISIONS NEVER TOUCHED THE READING PIPELINE — SAME GUARANTEE AS FIG. 15a.
BAR = CROSS-VALIDATED AUC AT TELLING ACCEPT FROM REJECT · 0.50 = UNREADABLE · 1.00 = FULLY DETERMINED
Fig. 15b
How to read
every decided, non-withdrawn paper of 2018–2026 (39,484; 2019 carries no decision records in this corpus), each reduced to what its reviewers criticised before any rebuttal: the tally model sees only fourteen numbers — how many negative units each of the twelve objects of scrutiny drew, plus the paper’s total praise and hedged counts — and a logistic regression guesses the decision under five-fold cross-validation. If decisions were unreadable from criticism, every bar would sit at 0.50; the review-count-only bar (0.51) is that floor made visible — volume alone carries nothing. The scores bar is a reference, not a rival: it answers “how far is anything text-side from what the panel’s own numbers already say.”
Result
— real, and far from everything (adding review count moves it only to 0.652; ). The heaviest count, once again, is novelty (−0.36 per unit, with problem framing −0.22 and baselines −0.17 behind it) — the third independent appearance of the novelty premium, after Plate VI’s score weights and Plate XIV’s tariff — while compute cost and robustness carry ≈0 weight, echoing their empty gaps in Fig. 15a. And one drift the pooled bars hide: read year by year, while the text ceiling gives up almost nothing (0.73 → 0.71) — the verdict has drifted away from which topics were criticised, far more than from the criticism’s language.
What was done
target declared before results were seen (accept = the decision string contains “accept”; withdrawn and desk-rejected excluded); features are counts of negative units by object, from initial reviews only; text ceiling = mean-pooled local embedding of the same units’ judgment sentences; stratified 5-fold CV, seed 7; base acceptance rate 39.5%. Script: build_decision_commit.py, output shipped as decision-data.json.
Reading it
association, not mechanism: reviewers may criticise novelty because they are rejecting, and an AUC cannot tell the direction. Per-year AUCs are computed within their own year, so the moving score-ruler does not contaminate them — but the pooled scores-reference does pool four scoring forms. And the fall is not unique to criticism: the scores’ own legibility slips too, though later and less (0.96 through 2023, then 0.94, 0.94, 0.89) — part of what fell is the decision’s readability from anything. Within that, the tally-vs-text contrast is the criticism-specific part, and the nine-year fall of the tally bar is a fact about legibility, not about rigor — it says the decision increasingly turns on something other than the topic-mix of criticism: the words, the rebuttal, or what no text records.
ACT V
The Encounter
when people meet, does judgment move · plates XVI–XXI
Overture · The Season
Corpus: ICLR 2026 · every public event, timestamped to the day · Sep 2025 – Mar 2026

One review season, replayed

Before asking whether judgment moves when people meet, watch the calendar the meetings happen on. Every public act of the ICLR 2026 cycle, on its real calendar day: 19,814 submissions crest at the deadline, 75,859 reviews land in a two-day burst, 180,146 discussion comments answer them, 5,216 papers withdraw — most in the week after the reviews arrive — and on a single day in January, 14,175 decisions fall at once. This overture replays a single season; the plates that follow read nine years of such seasons.

The withdrawal ribbon rises within days of the reviews becoming readable — and long before any decision exists.
Mar 08
BAND HEIGHT = REVIEWS SUBMITTED THAT DAY, √-SCALED TO EACH BAND’S OWN PEAK
The Season
How to read
each ribbon is one kind of public event, day by day (square-root vertical scale, so the deadline spike does not flatten everything else); press replay and the season unfolds — the playhead carries the date, sparks mark days with over a thousand events, and the counters accumulate the running totals. The shape worth noticing: . What that timing shows is when withdrawals happen, not why — but it does place most of them after the reviews were readable and before any verdict existed.
Plate XVI · The Shapes of Talk
Corpus: 199,031 review threads, 2018–2026 · pure reply structure, no machine reading anywhere

Every conversation has a skeleton

Why this exists
every other plate depends on a language model having read the text. This one deliberately does not: a reply tree can be counted from timestamps and authorship alone, so whatever it shows is immune to any error in the reading. Beneath every review grows such a tree, and it can be read without reading a word: did the authors answer, did the reviewer ever speak again, did the exchange become a real volley? Classified by turns actually taken, the nine years tell a sharp story — the conversation climbed, unevenly but unmistakably, from 9% reviewer return in 2018 to its , then collapsed in 2026 — the year submissions grew 70% and the number of reviews grew 62%.
Fig. 16a — A census of skeletons: what became of each year's review threads
.
Fig. 16a
How to read
every official review roots a thread; its shape is decided by the turns beneath it — a stump received no reply at all, a monologue was answered by the authors with the reviewer never speaking again, a return saw the reviewer come back once, a volley twice or more; the rare aside (replies from neither authors nor reviewer, 0.05–2% of a year) is the hairline at the top of the strip. Roles for 2018–21 are recovered from signature formats, so the earliest years are the least certain. The strip stacks each year’s threads to 100%. The brass layers — return and volley, the threads the reviewer came back to — rise from the baseline, so their share reads directly as a height; the dark stumps hang from the top, so silence reads as a thickness pressing down; monologue fills the middle. The numbers along the top are the came-back share, year by year.
Result
the reviewer came back to 43.8% of threads in 2025 — the high-water mark — and to 17.7% in 2026, while stumps grew from 9% in 2021 to 26%: one review thread in four now ends in silence before it begins. One care for the year-to-year line: each year’s discussion rules differ — response windows, mandatory acknowledgements — so the levels track the venue’s process as well as its people. In 2026, threads where the reviewer returned were accepted at 50–51% against 39% for monologues and 1.2% for stumps — read the direction both ways, since authors triage which battles to fight and which papers to quietly abandon.
Fig. 16b — Five real skeletons, pressed like botanical specimens
Five real discussions pressed flat: .
19A IS THE EVIDENCE — THESE ARE ITS SPECIMENS — the census above counts all 199,031 threads; this figure only puts faces to its shapes. Five discussions picked deliberately for their extremes (the selection rule is printed on each card), so read them as portraits of what a volley, a monologue, and a stump actually look like — never as a sample. What is drawn is raw record: who replied, and when.
Fig. 16b
What was done
five real 2026 discussions, chosen to show the range rather than to sample it: three with the most reviewer returns among threads of six to fourteen notes, plus the most monologue-dominated and the most stump-dominated discussion with at least three reviews. They are illustrations, not evidence.
How to read
actual 2026 discussions, each drawn in full: squares are official reviews, dots the replies beneath them — deep ink for authors, brass for reviewers, pale sepia for chairs, grey for the public. The left specimens hold long alternating vines — reviewer and author trading turns five and six deep; the right ones are all stumps and short monologues. Hover any specimen for its decision.
Interlude · One Event, Three Denominators

One event, counted three ways

The next three plates watch a single event — after the author’s rebuttal, did the reviewer’s written judgment soften? — and differ only in what they divide it by. The Rebuttal counts per unit written after the response: of what got written again, what moved. The Moves counts per author–reviewer exchange, joined to what the author did: which moves travel with movement. The Fate of an Objection counts per objection originally raised — the only denominator that can see the objections never spoken of again, and they turn out to be the large majority. Same event, three denominators; the denominator decides what each number can mean.

THE REBUTTAL ↓
Of the judgments written again after the response — how many soften, harden, or hold?
denominator: 130,650 post-response units
sees only what was written again
THE MOVES ↓
Which authorial move — new experiments, concession, contest — travels with a softened judgment?
denominator: 67,290 author–reviewer exchanges
sees what the author did
THE FATE OF AN OBJECTION ↓
Of every objection raised — what ever happens to it at all?
denominator: 551,595 objections raised
the only window that can see the silence
Plate XVII · The Rebuttal
Corpus: ICLR 2018–2026 · Fig. 17c: 2026 only

Author responses barely move the written judgment

130,650 units sit after the author response (2018–2026) — written by the reviewers who spoke again at all; most never do (Plate XVI). When a written judgment moves, it moves toward reinforcement 2.4 times more often than concession — and every kind of response, fresh experiments included, reinforces at least twice as often as it softens. Two limits, stated up front: this plate reads what reviewers wrote, so a score quietly raised without a word is invisible — the corpus keeps no rating history — and every recorded fate is a model's reading of the text.

Fig. 17a — Fate of a judgment after the author responds
After the authors respond, most written judgments simply stand.
COUNTING UNIT: THE POST-RESPONSE JUDGMENT — all 130,650 units reviewers wrote after a rebuttal, each carrying a recorded fate. Only reviewers who wrote again are counted — most never return (Plate XVI) — and a score changed without a word of text is invisible here. Plate XIX will re-count the same event per objection raised; the denominators differ, the story should not.
Fig. 17b — What the response contained, and what happened
Even fresh experiments mostly strengthen the original judgment.
strengthenedclarified onlyweakenedreversedeach bar = that trigger’s revisited judgments
Fig. 17c — Who yields? Post-rebuttal movement by position in a split verdict (ICLR 2026)
When ratings split by four points or more, the reviewer holding the lowest score is the one who yields — .
Fig. 17
What was done
when a reviewer writes again after the authors respond, the model that read the discussion records, for each judgment, whether it was strengthened, weakened, reversed, merely clarified, or left unchanged, together with what the reviewer names as the trigger. These 130,650 post-response units are what the three panels count. 17a: every post-response unit, flowing from the response into its recorded fate — each band’s width is the exact share of units in that state; . 17b: among units whose judgment visibly moved, split by a keyword reading of the recorded trigger (two rare triggers — unstated, and bare commitments — moved 163 judgments between them and are omitted); red = weakened or reversed, brass = strengthened, grey = clarified only. . Softening rates barely differ by topic (3.9–5.4% across all twelve objects).
How to read 17c
reviewers on 2026 papers whose ratings split by 4 points or more (two steps on a scale that runs 2, 4, 6, 8, 10), grouped by their own position; oxblood dot = share whose judgment weakened or reversed, brass dot = share that strengthened. The tinted band spans the yield rates of every group except the low scorer — only the low scorer’s ringed dot stands outside it. — yet even for them reinforcement beats concession, and unanimous panels barely stir at all (9,329 split panels; a reviewer's units are matched to their rating through their anonymous OpenReview identity, which succeeds for 95% of reviews). Two cares close the plate: every fate here is the model’s reading of prose — a courteous "I thank the authors and maintain my score" reads as reinforcement — and a rating changed without accompanying text never enters these counts; the corpus stores no rating history, so the score’s own needle is beyond this plate’s reach.
Plate XVIII · The Moves
Corpus: ICLR 2018–2026 · 67,290 author–reviewer exchanges with observable outcomes

What authors do, and what actually moves

Every author reply in a reviewer’s thread, scanned for seven quotable moves — the concession, the contest, the delivered experiment, the promissory note — and joined to whether that reviewer’s post-response units record any weakened or reversed judgment. An exchange exists only where the reviewer wrote again at all — most never do (Plate XVI). Across all 67,290 exchanges, — that is the number every move below has to beat.

Fig. 18a — The currency is effort
— no single move tracks softening as strongly as sheer volume of engagement.
Fig. 18b — The seven moves, effort held fixed
. “We respectfully disagree” keeps a +1.4 to +2.8pp softening premium in every reply-length stratum; every other move’s premium touches or crosses zero.
Fig. 18
Where to look
right panel: the dashed line is the null — a move that changed nothing puts every mark on it. Six of the seven bands touch or cross that line; only the contest’s clears it.
How to read
left, the share of exchanges in which the reviewer recorded a weakened or reversed judgment, by decile of the author’s total reply length (median words on the axis) — an association, not a payoff: longer replies also happen where there is more to fix, and where the reviewer was engaged enough to answer at all. Right, each move’s four filled dots are its softening premium computed inside reply-length quartiles, so a move common in long replies cannot borrow the effort effect; the hollow dot is the raw gap before that control, and the band spans the four strata.
Result
only the contest — detected in 5% of exchanges, though the seven detectors differ in breadth, so prevalences compare only loosely — stays positive in every stratum: +1.6 to +2.8pp in the three quartiles with enough contests to read, and +1.4pp in the shortest-reply quartile, drawn faint because it holds only 265 contests, under the 300-exchange floor the other strata clear. The rest flatten once effort is held fixed: the concession and the ritual thanks swing between −1.4 and +1.4pp with no consistent sign, and the promissory note keeps a faint non-negative residue (0 to +1.6pp). : within ±0.7pp of the line, negative in three of four strata, and the most negative average premium of the seven — though the ritual thanks and the plain revision also average fractionally below the line. An object-level check finds the same flatness at the point of impact: among exchanges charged on a given object, delivering experiments shifts softening on that object by less than half a point in either direction across all eight most-charged objects, with no consistent sign (novelty 1.1% with vs 1.3% without; empirical scope 1.7% vs 1.2%). Read the contest with care: authors choose to contest when they hold a strong hand, so this is the association of a chosen move, not the payoff of a tactic anyone could copy. And the outcome measure is conservative — it counts only judgment changes the reviewer put in writing; a reviewer quietly raising a score without re-engaging the argument is invisible here (Plate XVII reads the same limit from the other side). An embedding clustering of 722,406 reply sentences — a 26,428-exchange subsample of this universe, capped at 30 sentences per exchange (method § 10) — separates topics more than tactics and finds the same null.
Plate XIX · The Fate of an Objection
Corpus: ICLR 2018–2026 ·

Most objections are never spoken of again

Follow every objection — a reviewer's initial negative unit on one object — through the discussion phase: did the same reviewer return to that object after the authors responded, and if so, did the judgment soften or harden? The dominant fate, for every kind of objection, is silence.

Fig. 19a — Raised → revisited → softened / hardened, per object
. Of the 8% that are, most stand unchanged — and when the judgment does move, it hardens 2.5× more often than it softens.
COUNTING UNIT: THE OBJECTION RAISED — every pre-rebuttal objection, asking whether it was ever revisited; Plate XVII counted the same event per post-response unit. Broader denominator, same verdict.
revisited after the responseof revisited: softenedof revisited: hardened
Fig. 19a
Where to look
the top strip is the whole story: of all 551,595 objections, 92.1% occupy the near-empty left span — the same reviewer never mentions them again — with 6.5% revisited but unmoved, 1.0% hardened and 0.4% softened crowded at the right edge.
How to read
below the strip, each row is one object of scrutiny on a ×7-magnified axis; the grey bar is the share of that object's objections the reviewer explicitly revisited (4–12%); within it, verdigris marks objections that softened and sienna those that hardened.
Result
. The extremes are instructive: — precedent, once cited, is simply restated.
Fig. 19b — The interrogative: does an objection asked fare differently from one asserted?
Asked or asserted, an objection meets the same fate — revisited at 5.1% vs 5.2%, a null result reported as one. The only thing that moves the needle is raising it in both fields (7.4%): intensity, not grammar.
“ASKED” IS PLATE V's MEASURE — the share of a subject's objections phrased as questions, introduced beside the bill of faults; here the question mark is followed to its outcome.
Fig. 19b
A null result, reported as one.
How to read
each row is one fate — revisited, softened, hardened — and each dot is the share of objections that met it: brass for objections written as questions (in the review's [questions] field), grey for objections asserted (in [weaknesses]), violet for objections the reviewer wrote in both fields. Every row runs on its own 0→max scale. If the question mark changed an objection's fate, the brass and grey dots would sit apart.
Result
they coincide on all three rows: revisited 5.1% against 5.2%, softened 0.2% against 0.3%, hardened 0.6% in both — the interrogative is a register of politeness, not a different epistemic act. Only the violet dot steps away: — what moves the needle is intensity, not grammar.
Why ask
objections differ sharply in form — robustness concerns are posed as questions 29% of the time, novelty verdicts only 9% — so if form mattered, the question mark would be doing invisible work all over this atlas.
The count
109,062 reviewer-object objections from 2026 were anchored to the review's form fields; 5,000 anchored in neither field (summaries, strengths, limitations sections) are set aside, leaving 104,062 classified — 13,568 asked, 85,343 asserted, 5,151 in both — each joined to its post-response fate.
Plate XX · The Panel
Corpus: ICLR 2018–2026 · 1,009,592 units

Dissent, and the kinder judge

Reviewers of the same paper, read as a panel — four questions in one plate. Where do co-reviewers take opposite stances (Fig 20a)? How does the meta-reviewer’s tone differ from the panel they summarize (the three numbers beside it)? When the final decision disagrees with the panel’s lean, which way does it err (Fig 20b)? And how much of the ground does a panel actually walk (Fig 20c)?

Fig. 20a — Contested ground: share of papers where reviewers took opposite stances
Panels split most over whether the question matters at all.
Fig. 20a
How to read
each row is one object of scrutiny; the bar is the share of its papers on which two official reviewers landed on opposite valences (a paper counts only when at least two reviewers touched the object; bars drawn relative to the 38.6% maximum).
Result
.
The rail
the three numbers beside the figure are a separate reading with no figure of their own: across nine years of units, meta-reviewers run 13.7pp less negative and 10.5pp more positive than the panels they summarize — the judge is kinder than the prosecution. Part of that gap is genre: a meta-review summarizes and justifies a decision rather than inspecting a paper, so its milder mix is not by itself proof of a milder mind.
Fig. 20b — The bench: when the decision disagrees with the panel's lean
When the decision departs from its panel, it departs toward mercy: .
The overruled — public exemplars
Fig. 20b
What was done
for each of the 13,700 decided 2026 papers with three or more rated reviews, the panel's lean is its mean rating, and the accept line is the mean-rating cut that reproduces the venue's actual acceptance rate (5.0 in 2026). A paper counts as overruled when its decision lands on the far side of that line by at least half a point — rejected despite a mean of 5.5 or more, accepted despite a mean of 4.5 or less.
How to read
the curve is how often the decision agrees with the panel's lean, plotted against distance from the accept line, so agreement near zero is the interesting part; the figures beside it count the overruled cases in each direction, and the mirrored wings beneath draw those two counts against the shared accept line — a bench that erred with no direction would draw two equal wings.
Result
— when the decision departs from its panel, it does so toward acceptance about 2.7 times as often. The half-point margin is a choice, so the ratio was recomputed at margins of 0.25, 0.75 and a full point: it runs 2.4×, 3.9× and 3.9× — the direction survives any reasonable margin.
Reading it
the accept line is inferred from the outcomes, not published by the venue, so a paper counted as overruled may simply have been decided on information the mean does not carry — a chair's reading of the discussion, a policy on scope, a reviewer discounted.
Fig. 20c — The search party: what a panel covers, duplicates, and never sees
The panel duplicates some searches and leaves a third of the ground unwalked.
Fig. 20c
How to read
treat the panel as a search party over the twelve objects of scrutiny. For each object, the bar splits the 18,074 rated 2026 panels of three or four by how many of their reviewers examined it at all — from the void of no one to the bright band of three or more. The party is uncoordinated: the average panel covers 7.95 of the twelve objects, agrees on examining only 0.83 of them, and leaves 4.05 entirely unvisited. Duplication and blindness sit side by side — , and the objections that decide papers are among the least searched: . Coverage here means one unit of any valence — a low bar, deliberately.
Plate XXI · The Deliberation
Corpus: 17,848 ICLR 2026 panels (21a) · fourteen deliberations scored in full (21b) · 49,441 panels 2018–2026 (21c–d)

Watch a panel think

First, how the 17,848 panels of ICLR 2026 usually end — then fourteen real deliberations, written out like music. Each reviewer is a staff and each unit a note, placed in the order it was reasoned; the dashed barline is where the authors respond, and what follows it carries the fate of each judgment — hardened, softened, or reversed. Faint verticals bind reviewers who touched the same object: ink where they agree, sienna where the same object drew opposite verdicts. Hover any note to read the unit.

Fig. 21a — How deliberations usually go
Six panels in seven — 85.7% — never visibly move at all.
Fig. 21a
How to read
one bar per ending — silence, procedural, softening, entrenchment, contested — its length the share of panels showing it.
What was done
all 17,848 rated 2026 panels with two or more reviewers whose units could be matched to their rating, classified by the field in which the model records whether a reviewer's judgment strengthened, weakened, or reversed after the authors replied. The five endings overlap (a split panel can also soften) and the levels are machine-read, so treat them as a census of what is visible in writing: movement of any kind is rare — — and when judgments do move after the response, .
Fig. 21b — The case files: fourteen deliberations, scored in full

Everything else in this atlas averages thousands of panels; this figure does the opposite. Below are individual papers — fourteen real 2026 discussions, hand-picked as extreme, legible specimens of the five endings above (not a random sample; the pull-quotes on each card are the panel's final ratings). Open a case to read its full score.

negative unit positive conditional uncertain ▲ hardened after response ▼ softened ↺ reversed
Fig. 21b
How to read
rows are the paper's reviewers, ordered by rating (badge at left); notes are their logic units in reasoning order, colored by verdict; the brass barline marks the author response, and the notes after it are the same reviewer returning — the triangle above a note records whether that judgment hardened or softened, the rare ↺ a reversal. The vertical threads are the panel's harmony and dissonance: the same object of scrutiny judged alike (ink) or oppositely (sienna) by different reviewers. Fourteen scenarios were chosen to span the repertoire — softenings, entrenchments, reversals, ten-point splits, unanimity — and every one links to its public page on OpenReview.
Fig. 21c — The repertoire: the five endings and their verdicts
— and co-reviewers barely inspect the same things strangers would.
HOW A DISCUSSION ENDS × HOW OFTEN THE PAPER IS ACCEPTED · 2026
DO CO-REVIEWERS EVEN INSPECT THE SAME THINGS?
Fig. 21c
How to read · left
each row is one ending of a panel's discussion phase (silence: no post-response units; procedural: units but no recorded movement); the bubble sits at the share of that ending's decided 2026 panels that were accepted, its area is the ending's share of 2026 panels, and the dashed rule is the venue's own 2026 rate (39%) — bubbles slide out from that rule, so a fate that made no difference would never leave it. In four panels out of five, no judgment moves at all; the nine-year composition, year by year, is Fig. 21d below. The acceptance rates are association only: entrenchment accepts most because reinforcement includes reviewers reinforcing praise, and unanimity of silence often means unanimity about rejection. (right) — reviewers of one paper attend to only marginally more of the same things than strangers do. So the disagreement measured elsewhere in this atlas begins before any verdict is reached, at the choice of what to look at. And the dissonant objects — where co-reviewers reached opposite verdicts — are revisited after the response slightly less often than the agreed ones: the chords do not resolve.
Fig. 21d — The great quieting: 49,441 panels, year by year
Conversation bulges in 2022 and is squeezed shut by 2026 — silence falls 51% → 20%, then climbs all the way back to 52%.
Fig. 21d
Where to look
the shape is the finding: the inked strata rising from the floor are the panels whose discussion phase did anything at all — judgments moving (verdigris, violet, sienna) under talk that moved nothing (dim grey) — and the unmarked page above them is silence. The strata swell toward 2022, the most conversational year the venue has had (judgments moved in 29% of panels), then are squeezed to 13% by 2026, as submissions grew past nineteen thousand; the figures on the boundary are silence's share, .
How to read
each year's column stacks to 100% of that year's panels; same categories as 21c, computed per year over all 49,441 panels, 2018–2026. On first view the strata rise out of a flat 2018 and settle into the bulge.
Reading it
descriptively only: the format, the volume, and the culture all changed together, and this figure cannot separate them.
ACT VI
The Higher Court
the chair’s court — what it weighs, whose words it borrows · plates XXII–XXIII
Plate XXII · The Higher Court
Corpus: 56,736 meta-reviewer units (90% from 2024–26) · 6,547 borderline 2026 panels · association, not cause

The judge above the judges

Above every panel sits an area chair who writes a meta-review — and those were read the same way: . The higher court has its own grammar: it does not re-litigate the evidence, it weighs the verdicts. And at the margin, where panels sit just below the accept line, its choices reveal which objections it declines to forgive.

Fig. 22a — Two grammars: what the panel inspects, what the chair weighs
The chair does not re-litigate evidence — it weighs verdicts: statistics, positioning, novelty.
Fig. 22b — What the bench forgives at the margin
At the margin, only a clarity charge reliably blocks the pardon.
Fig. 22a
How to read
each line is one object of scrutiny; its left end is the object's share of all 952,856 official-reviewer units, its right end its share of the 56,736 meta-reviewer units (nine-tenths of them from 2024–26). The chair's attention is not the panel's: — the objects a chair can rule on from the reviews alone — while . Meta-review units are also positive 30% of the time against the reviewers' 19%: the higher court speaks in verdict-language, not evidence-language.
Fig. 22b
How to read
among the 6,547 borderline 2026 panels (mean rating 3.5–5.0, decided), each point compares how often lifted papers carried at least one negative unit on the object against rejected papers at the same mean rating (matched inside 0.25-point bins; whiskers are a 95% interval from resampling whole papers). Left of zero, the charge travels with rejection. The tilts are small and only — but the ordering echoes the tariff of Plate XIV: , while . Clarity is the exception worth naming: it costs a reviewer nothing in score, yet at the bench it blocks the pardon.
Fig. 22c — The redemption: what the chair praises when it lifts a paper
75% of a lifted paper's meta-review is praise — of its numbers and its positioning; on rejected borderline papers, praise falls to 8%.

Fig. 22c — Share of all meta-review units on those same borderline papers that praise each object, for lifted papers against rejected ones. On a lifted paper roughly three-quarters of the meta-review is praise, and its favorite subjects are the numbers and the positioning — "the improvements are significant, the comparisons honest" — while on a rejected borderline paper praise of any kind nearly vanishes. Read as the language of justification, not independent evidence: the meta-review is written by the person who has already decided.

Plate XXIII · The Borrowed Verdict
Corpus: ICLR 2018–2026 · 8,439 papers · every meta-review unit matched to the reviewer unit it most resembles

Whose words reach the decision

Match every meta-review unit to its nearest reviewer unit on the same paper — nearest in wording, not in time — across the 8,439 papers of 2018–2026 with a meta-review of three or more units and at least two official reviews. Three tiers fall out: near-copies, echoes, and the area chair’s own words — and the borrowing has a direction.

Fig. 23a — What the meta-review is made of
. But criticism is borrowed more than praise — at every cutoff.
PLATE XXII COMPARED WHAT THE CHAIR WEIGHS — this plate asks where the chair's words come from: matched to the panel's, tier by tier.
Fig. 23b — When the AC borrows, from whom
Echoes concentrate on one reviewer (60% from a single voice), and in rejections they lean to the harsher side of the panel, 60/40.
Fig. 23
Where to look
the bars of 23a left-align the borrowed mass, so the borrowed edges compare directly: , and the gap survives any cutoff (+7.0 / +6.6 / +3.3 at ≥ 0.70 / 0.75 / 0.80 — thresholds are conventions, so the comparison was recomputed at all three). In 23b the dashed line is blind borrowing: rejections escape it to the harsh side, 60/40; accepts do not (51/49, interval crossing even).
How to read
tiers by cosine similarity — a 0-to-1 score for how alike two pieces of text are once embedded — between the meta unit’s reasoning and its nearest reviewer unit: near-copy ≥ 0.90, echo 0.75–0.90, below that the AC’s own words. The decision is written, not compiled — — but what is borrowed tilts: negative meta units echo reviewers at 28.8% against 22.4% for positive ones, and on papers with three or more echoed units, 60% of echoes come from a single reviewer.
The lean
among panels where every reviewer’s rating is known and the panel is not unanimous, — a lean, not a rule; in half of panels the echoed reviewer’s rating ties an extreme, the coarse scale allowing no verdict. Nearest-neighbor attribution measures textual proximity, not causal influence: an AC and a reviewer may both echo the paper itself.
The ledger
8,439 papers of 2018–2026 with ≥ 3 meta-review units and ≥ 2 official reviews; 31,598 meta units matched; the who/side analyses use the 3,334 panels with full, non-unanimous rating panels; the side test keeps the 2,220 decided rejections and 752 accepts where the echo-weighted rating differs from the panel mean at all.
ACT VII
The Measure
can a rating be trusted as a measurement · plates XXIV–XXV
Overture · The Tide
Corpus: ICLR 2026 & 2025 · every official review, binned by days before the filing peak

The early review is a different document

Reviews filed weeks early run harsher and longer than deadline-day reviews — in both years.
The Tide
Why this is here
this act asks whether a rating can be trusted as a measurement, and before its plates listen to what a score sounds like, one property of reviews as documents is worth having in hand: they are not interchangeable draws. When a review was written already goes with how harsh and how long it is, and no plate in this atlas controls for filing time — so this stands as a caveat over the whole atlas, backward as well as forward. It is filed in this act because filing time is a fact about the calendar rather than about the text — a confounder of the measurement, not part of what reviews say. The plates where a score first entered the story (VI, XIV, XV) each carry a pointer back to this one, and the reader’s note names it before Act I.
What was done
every official review of ICLR 2026 and 2025 was placed in a bin by how many days before that year's filing peak — the de-facto deadline — it was first posted, and each bin's mean rating, median length and mean number of units were computed.
How to read
the horizontal axis runs from two weeks early, on the left, to after the deadline, on the right. Left: mean rating, shown as the deviation from the same year's deadline-day mean so the two scale regimes never share an axis; whiskers are bootstrap 95% intervals. Right: median length in words. The gradient is monotone and replicates in both years: , and carry more units of logic (5.68 vs 5.40). Read without a causal story — early filing is not randomly assigned: reviewers who finish early may simply be those with the most to say, and a paper that invites a swift, thorough rejection gets one. What this establishes, robustly, is only a correlation: when a review was filed goes together with how harsh and how long it is. ⚖ the scale itself changed four times over these years — method § 10 · the moving ruler
Plate XXIV · The Ladder
Corpus: ICLR 2026 · official ratings joined at display time

What each score sounds like

Why this exists
a question the researcher set down before any number existed: why is a low score low, and why is a high score high? Ratings and units were extracted independently, so joining them can answer it. Some of what follows is near-tautological (a low score has more negative units); the caption marks which findings are not.
What was done
join every 2026 review's official rating — never shown to the extraction pipeline — to its logic units: what a 2 argues about, what an 8 still checks, and where on the ladder the reviewer's own confidence lives.
Fig. 24a — The ladder: verdict mix per rating (ICLR 2026, whose ratings run 0–10 in steps of two)
The mix, not just the amount, is the score.
PLATE VI’S JOIN, READ BACKWARD — the Early Verdict asked how well the verdict mix predicts the score; this ladder reads the same join the other way: what each score is made of.
Fig. 24b — What a 2 argues vs what an 8 argues (share of units per standard)
The standards swap with the verdict: .
rating 2rating 8
Fig. 24c — The relocation of scrutiny: share of reviews that inspect each object, per rating
Scrutiny relocates rather than shrinks: — .
Fig. 24d — Same paper, split verdict: share of papers where each side condemns the object
On the same paper, both sides inspect theory and method at identical rates and split only on the verdict —
PLATE XXI ASKED THIS OF EVERY PANEL — contested ground counted opposite stances across all papers; here only the pairs already split on the score, asking what each side condemns.
LOW SIDE
HIGH SIDE
low condemns ← left wingright wing → high condemnspale wing = inspected at all (any stance)wing length = share of that side condemning the object
Fig. 24e — The spread: how much co-reviewers of one paper disagree, and whether it matters
93% of multi-review papers disagree, half by four points or more — yet at a fixed exact mean, unanimous and divided panels are decided alike (92% vs 92% at mean 6.0; 12% vs 13% at mean 4.0).
Fig. 24
How to read 24a
each row is one rating; the bar is the verdict mix of its units; the thin brass underbar is the relative number of reviews at that rating; the right column reads n reviews · units per review · mean reviewer confidence (1–5). Ratings are joined at display time only; extraction never saw them. Five findings that are not the tautology (low score = more negative): (1) Scrutiny relocates rather than shrinks — novelty largely disappears from an 8's review (a novelty unit appears in 20% of 2-reviews but only 9% of 8-reviews) and baselines (42% → 29%), but inspects theory, method design, and compute more than a 2 does. (2) A 2's praise leans consolation — both courts of praise lead with empirical scope, but a 2's rare positive units tilt toward the packaging (clarity takes 13% of them against an 8's 8%, framing 8% against 6% — "well written, nice motivation") while an 8's run deeper into theory (16% vs 13%). (3) There are many ways to fail — within-rating diversity of what reviews discuss falls monotonically as the rating rises, so that reviews rating a paper 0 are the most varied in subject and those rating it 10 the most alike: unhappiness is diverse, excellence has one shape. (4) In split verdicts on the very same paper, both reviewers inspect theory and method at identical rates and merely disagree on the verdict — but novelty is a weapon only the low scorer draws (condemned by 21% of 2-givers vs 4% of their 8-giving co-reviewers, across the 1,789 papers rated both 2 and 8). Quieter still: the reviewer's own confidence follows a U (4.2 at 0 · 3.4 at 6 · 3.9 at 10), and mid-scale criticism is the most actionable (78% of a 2–4's units carry a fix vs 56% at 10). (5) The decision is a mean, not a variance — 93% of multi-review papers carry disagreement (half of them by 4 points or more), and disagreement peaks for mid-scoring papers; yet holding the exact mean fixed, a unanimous panel and a divided one are accepted at the same rate — 92.1% vs 91.8% at mean 6.0, and 12% vs 13% at mean 4.0, where the threshold actually bites. A naive banded version of this analysis suggests "consensus is fate" — that is a mean-composition confound, and the panel averages it away.
Plate XXV · The Measurement
Corpus: ICLR 2018–2026 · all rated official reviews · a property of the system, not of any reviewer

How repeatable is a score?

Treat a paper's rating as a measurement and ask what any instrument must answer: if the measurement were repeated — same paper, another qualified reviewer — how much would it move? The NeurIPS consistency experiments answered by having two independent committees review the same papers; this is the observational shadow of that experiment, computed from every co-review of the same paper across nine years of ICLR.

Fig. 25a — The paper's share of the score: how much of a rating is the paper, not the reviewer (by year)
In every year, most of a score's variance is not the paper — its share holds at 0.34–0.45 ().
Fig. 25b — Same paper, another reviewer: the score is a distribution
Same paper, another reviewer: the score is a distribution — draw from it.
The other reviewers of a paper average
EACH BAR = SHARE OF NEXT-REVIEWER SCORES, GIVEN THIS PANEL MEAN
Fig. 25
How to read 25a
the intraclass correlation is the share of rating variance attributable to the paper rather than to the reviewer draw; — in every year, the majority of score variance is not the paper. , but its coarser even-numbered scale mechanically deflates the estimate, so the cross-year drop should not be over-read — and 2020's four-point form carries a milder dose of the same coarseness, so its true share may sit somewhat above its dot. 25b: each bar is the real ICLR 2026 distribution of one reviewer's score conditioned on what the paper's other reviewers averaged; the draw button samples from exactly that distribution. Framed deliberately as measurement noise: this is a property of the system under its constraints — anonymous volunteers, four reviews, deadline pressure — not a verdict on any reviewer. ⚖ the scale itself changed four times over these years — method § 10 · the moving ruler
Fig. 25c — Your split, in context
Measure your own panel's spread — a wide split is not a death sentence: panels split by 8 points accept 47%, unanimous ones 30%.
Enter a panel's scores
Fig. 25c
How to read
type any panel's scores and they land on the upper rule; a pair of dividers measures the spread (max − min), then transfers that measurement onto the lower axis, where every real ICLR 2026 panel with three or more rated reviews is stacked by its spread. Bar heights are √-scaled so the rare wide splits stay visible; the green fraction of each bar, and the percentage above it, is the share of those panels accepted. One quiet regularity worth knowing: a wide split is not a death sentence — panels split by eight points were accepted 47% of the time (a thin bin — 168 panels), while perfect unanimity was accepted 30% of the time (1,319 panels), because unanimity is most often unanimity about rejection. The spread of scores has also not collapsed over the decade — the entropy of each year's score distribution — a 0-to-1 measure of how evenly the scores spread across the available levels, where 1 would mean every level used equally — normalized to that year's own scale so that changes of scoring form do not masquerade as changes of behavior, holds between 0.74 and 0.82 in every year but one (2020's four-point form reads 0.92) — so "everyone scores a six now" is not what the record shows. ⚖ the scale itself changed four times over these years — method § 10 · the moving ruler
Fig. 25d — The counterfactual conference: redraw every panel, count the flips
Redraw the panels and most borderline decisions come out as a coin toss — 58% of accepted papers, 37% of all decisions, flip on a fresh draw.
Fig. 25d
How to read
a deliberately simple model, stated in full. Each paper's underlying value is estimated from its observed panel mean by the variance decomposition of Fig. 25a (shrunk toward the venue mean in proportion to the reliability of a panel its size); a fresh panel of the same size then re-reads it, and the venue's own empirical acceptance curve — the observed P(accept | panel mean) — makes a fresh margin call. The line is the probability, over 4,000 such redraws, that the decision comes out differently, by observed panel mean. In the band where most decided papers live the redraw is close to a coin toss (), falling toward certainty only at the extremes. Under this model 37% of all 14,175 decided 2026 papers — and 58% of the accepted ones — would receive the other decision from a fresh draw; on 2025's finer scale the accepted-paper figure is 45%. The NeurIPS experiments, which ran real second committees, found roughly half of accepted papers un-reproduced. This model lands in the same range, which is reassuring but is not a validation: one is a real second committee, the other a simulation built from this atlas's own variance estimate. This measures the system's precision under its constraints, not any panel's diligence; and because the margin call is itself part of the redraw, it counts the decision layer's own latitude alongside the panel draw.
Fig. 25e — One alternate opening night
One press of the button, one alternate conference — never the same twice.

Fig. 25e — Every square is one of the 5,359 papers accepted at ICLR 2026, ordered by panel mean from the strongest accept down. Press the button and the model deals one alternate conference: brass squares survive the redraw, dark ones do not — and their seats would largely be taken by papers rejected this time (a quarter of rejections flip the other way). Each press is a new deal; no two alternate conferences are the same, which is the point.

ACT VIII
Eras & Territories
the instrument measured — is any of it particular to now, or to one field · plates XXVI–XXIX
Overture · The Field
Corpus: ICLR 2018–2026 · OpenReview public record · 52,460 papers

ICLR itself, 2018 → 2026

Before asking what belongs to an era or a territory, start with the venue itself, apart from any of the review analysis. These are plain counts from ICLR's public record on OpenReview: how many papers were submitted, how many were accepted, how long the reviews ran. , and that growth is the backdrop for everything in this atlas.

Review length fell for three straight years from its 2022 peak, then held flat; reviews per paper stepped from three to about four in 2021 and stayed.
Full census table
Census
How to read
one small chart per quantity — submissions, acceptance rate, reviews per paper, median review length — each plotted against the year, with the full table underneath. Counted directly from the OpenReview record; no machine reading is involved in this plate. Acceptance among decided, non-withdrawn papers; 2019 decisions are not recoverable from the public record and are shown as —. Withdrawn counts are only meaningful from 2025 (earlier years delisted withdrawals). The quiet story: — while the number of reviews a paper receives stepped up from three to about four in 2021 and has stayed there.
Plate XXVI · The Drift
Corpus: ICLR 2018–2026 · 1,009,592 units

Nine years of shifting scrutiny

Why this exists
the nine-year corpus was built for this question. Everything else on this page describes reviewing as it is; only a time series can say whether "as it is" is recent.
What was done
the same taxonomy, applied to 1,009,592 units from every ICLR review 2018–2026. Share of reviewer attention per year; the four biggest movers are named, the rest recede.
Fig. 26a — What reviewers inspect, 2018 → 2026
What reviewers inspect drifted over nine years, toward compute cost and statistical rigor.
THIS PLATE READS THE ENFORCEMENT, NOT THE LAW — what gets checked, and how often, drifts below. Plate XXVII holds a statute's text — novelty's stated definition — against the same nine years, and finds it unmoved.
2026
Fig. 26b — The standards they invoke, 2018 → 2026
The standards they invoke drifted with it.
PLATE II SHOWED THE COUPLINGS MOVING — Fig. 2b tracked which standard each object is judged by; this is the raw ingredient underneath, each standard's share of the year's units.
Fig. 26c — The hardening: nine-year vitals
Three vital signs of the venue, read across nine years.
Fig. 26
How to read 26c
three nine-year vital signs — how often a post-rebuttal judgment weakens or reverses, how much of a review is written after the rebuttal, and how many units a reviewer produces. The hygiene turn: ; . Among standards, while — scrutiny has drifted from what is this idea? toward what does the evidence cost, and does it cover the claim? Negativity in the nine-year corpus holds near 77% all nine years (the 2026 corpus reads 71.8% — a different memo style, the same story). The share of units the labeller places in the same category on a held-out check is the same in the early years as in the late ones (0.81), so the drift is a change in what reviewers wrote, not a change in how well the labels fit older text. The moves are small in points but far outside noise — a reviewer-clustered bootstrap (2,721 reviewers in 2018 against 73,427 in 2026) puts each of the four named moves at eight to sixteen times its standard error; in relative terms and compute cost's rose 67%. What the series cannot separate is reviewers changing from submissions changing: it measures what got written, not why. And the field is hardening: the share of post-response judgments that weaken or reverse has halved, 8.2% in 2018 to 4.2% in 2026 — though 2018 rests on only 1,239 post-response units.
Fig. 26d — The currents: four distributions, nine years — as a flow, or as terrain
Four distributions morphing through nine years — turn the dial, or stack the years like a mountain range.
2018
SCORES GIVEN
REVIEW LENGTH · WORDS
UNITS PER REVIEWER
THE SHAPE OF INFERENCE
BAR HEIGHT = SHARE OF THE YEAR’S REVIEWS · EACH PANEL SCALED TO ITS OWN TALLEST BAR
Plate XXVII · The Unamended Code
Corpus: ICLR 2018–2026 · 13,138 normative novelty sentences

Novelty's stated definitions look the same in 2018 as in 2026

The Drift plate showed enforcement changing — which objections are raised, how often, at what price. This plate asks about the statute itself. It was run expecting to find a change: nine years spanning the arrival of large language models seemed likely to have rewritten what counts as a new contribution. The result is negative, and is reported as a negative result. The test: fit the code of novelty on 2026 sentences alone, then hold every earlier year against it. If the law had been rewritten, 2018 should sit measurably farther from the 2026 code than 2026 does from itself.

Fig. 27a — Distance of each year’s sentences to the 2026 code
2018’s reasoning fits the 2026 rulebook almost as well as 2026’s own held-out half does — .
THE THIRD OF THREE NINE-YEAR VERDICTS — Fig. 2b found the coupling grammar still and Plate XXVI found attention moving; this plate holds the law’s own text still.
Fig. 27b — The clause mix, year by year
The same six clauses in every year — with one slow tilt: .
Fig. 27
How to read
left, the mean embedding distance of each year’s normative novelty sentences to the nearest clause centre — the average position of a cluster of similar sentences — in a code fitted on 2026 sentences only; the dashed line is the calibration — a code fitted on half of 2026 measured against the held-out half (0.152). Every year 2018–2026 sits within 0.008 of that line. Stated with its uncertainty: — and sentences cluster within reviews, so even that overstates the certainty), so the null is a claim about size, not undetectability — a drift of a few percent in the rule’s phrasing, beside the double-digit relative swings in its enforcement one plate back. The lower strip draws exactly this — each year’s deviation in SE units against a ±2SE corridor — and the bars beneath place the phrasing drift beside Plate XXVI’s four enforcement movers (different rulers, shown for scale, not as a test). An embedding distance can also only see changes large enough to shift the phrasing, and the six clauses are coarse. Right, each year’s sentences assigned to the 2026 clauses: all six clauses are present in every year of the record — the differentiation clause holds 36–50%, the combination clause 29–35%, the transfer clause 5–7%. One slow tilt is real: the differentiation clause’s share slips from 50% in 2018 to 38% in 2026, with the combination clause absorbing most of it — a drift in which clause reviewers reach for, not in what any clause says, so it belongs with Plate XXVI’s enforcement story rather than against this plate’s null. Two owned limits: sentence embeddings could miss a subtle semantic shift that preserves register; and the transfer clause rides on the code’s one genuinely mixed cluster (most of its sentences describe applying a known technique in a new setting, though some read equally well as ordinary differentiation claims), so its band is the least certain of the six. What changed across nine years is the docket’s traffic and its price (Plates XXVI, XIV) — not its text: the fashions of scrutiny moved; the stated standard of newness did not.
Plate XXVIII · The Oracle
Corpus: ICLR 2026 · 19,474 reviewed papers with a primary area · association, not destiny

Choose a field. Face its tribunal.

Why this exists
a comparison, not a claim: every paper carries a primary area, so the criticism profiles can be split by field without any further modelling. Nothing here was predicted, and field differences mix what a field submits with how it is judged.
How to use it
pick the primary area a paper would be submitted under, and the panel below recomputes from the real record: how often papers in that field drew at least one negative unit on each object of scrutiny, against the venue-wide baseline. The differences are the field's signature — what its reviewers reach for first.
THE CONSEQUENCE’S INDICATOR, SPLIT BY FIELD — the same “≥ 1 negative unit on the object” join that Plate XV read against decisions, here split by primary area against the venue-wide baseline.
BAR — SHARE OF THE FIELD’S PAPERS WITH ≥ 1 NEGATIVE UNIT ON THE OBJECT, 0–100% · INK TICK = VENUE-WIDE · ROWS SORTED BY DEVIATION
Plate XXVIII
How to read
each bar is the share of the field's papers that drew at least one negative unit on that object; the ink tick is the venue-wide baseline, and the right column is the deviation in points — sienna when the field is charged more than the venue, verdigris when less. Conditional frequencies only: fields differ in what they submit as much as in how they are judged.
Fig. 28 — The archipelago: all twenty-one primary areas, drawn as a sea chart
Fields criticized alike drift together — and each island's coastline is its fingerprint: .
Fig. 28
How to read
every island is one primary area, placed by multidimensional scaling — a method that puts similarly-treated items near each other on a map — of its criticism fingerprint — fields that are criticized in the same way drift together, and the theory-judged fields form their own archipelago in the east while the applied coasts lie west. Island size is the number of 2026 papers; the shape of the coastline is the fingerprint itself: each of the twelve objects of scrutiny is a fixed compass direction (the rose gives the bearings), and a cape grows toward every object the field is charged with more often than the venue baseline, a bay where it is charged less. So , and . Hover an island for its exact charges; the chart encodes association, not destiny.
Plate XXIX · The Watermark
Corpus: ICLR 2018–2026 · 199,031 official reviews · corpus-level traces only — no single review is accused

A watermark appears in 2024 — and fades while the practice spreads

Why this exists
since 2023, some unknown share of reviews has been written with a language model's help. No review discloses this, and per-review detectors are unreliable — they flag non-native English as machine text and are trivially evaded — so this plate never asks which review. It asks what a corpus-level instrument can honestly measure: whether the record as a whole moved toward the documented signature of machine-assisted prose. One published estimate already exists for one year — about 10.6% of ICLR 2024 review sentences were substantially LLM-modified (Liang et al., ICML 2024). This plate extends the question across all nine years with a simpler, fully reproducible instrument.
What was done
a thirteen-word signature list was frozen before any counting, taken verbatim from the two studies that quantified LLM-preferred vocabulary (Liang et al., ICML 2024; Kobak et al., Science Advances 2025): commendable, meticulous, meticulously, intricate, pivotal, versatile, delve, delves, delving, underscores, underscoring, showcases, showcasing. For every year, the share of reviews containing at least one of these words; the expectation from the 2018–2022 trend alone; and the excess above it. Alongside: how alike co-reviews of the same paper have become, and how the marked minority behaves. The full plan, hypotheses and decision rules were written down before computation (method § 12).
What this cannot say
nothing here attributes any trend to AI, and the watermark is ink, not a census. The list is a lower-bound tracer with a shelf life: newer models prefer different words, ICLR 2026 required reviewers to disclose LLM use, and a reviewer can delete the tells in one editing pass — so the 2026 fade is ambiguous three ways (different models, the new rule, or cleaner laundering). A fourth explanation was checked and excluded: . And the newest outside estimate points the other way: a January 2026 study applying the ICML likelihood method reads 26.7% of ICLR 2025 reviews as LLM-involved (Sharma et al., arXiv:2601.20920), rising while this tracer’s mark fell. The one year the record can calibrate against says the same: in 2024, when an independent method estimated one review sentence in ten machine-modified, this tracer marked only 7.8% of reviews — the watermark understates the practice even at its darkest.
Fig. 29a — The spike and the fade: the frozen signature, 2018 → 2026
The thirteen-word signature holds near one percent for six years, spikes to 7.8% of reviews in 2024 — and is back to 3.1% by 2026.
PLATE XXVII HELD THE LAW’S TEXT STILL — novelty’s stated standard reads the same in 2018 as in 2026. This plate finds the pen moving instead: the statute did not change; the diction did, in one sharp step.
SHARE OF THE YEAR’S REVIEWS CONTAINING ≥1 OF THE 13 FROZEN WORDS · DASHED = 2018–2022 TREND EXTRAPOLATED · ’24 = FIRST FULL POST-CHATGPT CYCLE · ’25 = OFFICIAL LLM-FEEDBACK PILOT · ’26 = DISCLOSURE REQUIRED
Fig. 29a
How to read
if the reviewer’s diction had merely continued its drift, the ink line would track the dashed counterfactual — and through 2023 it does, to within a third of a point (2023 sits at 1.2% against an expected 1.5%, ). Then the step: , . . A negative-control list of ten ordinary reviewer words (interesting, unclear, convincing…), run through the identical machinery, shows no spike — it drifts mildly below its trend as reviews shortened, which makes the marker spike conservative, not inflated. The excess is a floor, not a count: a review whose machine-assisted text happens to avoid all thirteen words is invisible here.
Fig. 29b — The tells rotate: each word against its own pre-LLM trend
The 2024 tells are not the 2026 tells: the delve family dies to below its human baseline while underscoring is still climbing.
A DETECTOR’S SHELF LIFE, MEASURED — any fixed word list ages with the models that made it famous; whatever replaces the 2024 vocabulary will not announce itself.
CELL = THE WORD’S FREQUENCY VS ITS OWN 2018–2022 TREND · WARM = ABOVE TREND · COOL = BELOW · ×N = FOLD · ROWS SORTED BY 2026 FOLD
Fig. 29b
How to read
each row is one frozen word; each cell holds its document frequency against its own extrapolated trend, so ×1.0 (pale) means “exactly as often as human drift predicts.” In 2024 every row but one runs warm — meticulously ×14.4, meticulous ×12.6, delves ×11.8, pivotal ×11.5, intricate ×10.6. By 2026 the column has split: — undershooting it, as if the word had been scrubbed — along with showcases (×0.51) and showcasing (×0.98), while underscoring (×253, from near-zero), underscores (×6.5), meticulous (×5.1) and commendable (×4.7) persist. One list, two fates: the vocabulary of the 2023-era models burned bright and burned out; a quieter residue is still spreading. For scale, the biggest 2026 vocabulary shifts of all are the field’s own subject matter — qwen appears in 9.2% of 2026 reviews, llama in 6.4%, deepseek in 2.6%, against essentially zero before — which is why the tracer is restricted to style words: topic drift is real, loud, and not evidence of who held the pen.
Fig. 29c — The convergence: co-reviews grow alike in words — and the object overlap turns out to be set size, not choice
Two reviews of the same paper are 14% more alike in wording under one fixed vocabulary; their apparent convergence in what they inspect dissolves against a size-matched null.
PLATE XXV READ THE NUMBER’S REPEATABILITY — how far two juries of the same paper agree in score. This panel reads the same pairs in words: whatever is happening to the diction, co-reviews are converging on each other, not on the corpus.
WORDING · TF–IDF COSINE
OBJECTS INSPECTED · JACCARD
INK = MEDIAN SIMILARITY OF CO-REVIEWS OF ONE PAPER · FAINT = RANDOM CROSS-PAPER PAIRS · VERTICAL RULES = REVIEW-FORM CHANGES · HOLLOW POINTS = FIXED SHARED VOCABULARY, 2024–26 · RIGHT, BRASS DASH = SIZE-MATCHED NULL (SAME SET SIZES, RANDOM CONTENTS)
Fig. 29c
How to read
left: for every paper with two or more reviews, the median pairwise TF–IDF cosine of its reviews (ink), against the median for random pairs from different papers (faint). If reviews were homogenizing wholesale, both lines would rise together; if the review form were doing the work, the line would jump at form changes (the vertical rules) and hold flat between them. Neither happens: the within-paper line climbs inside the constant-form window — 0.277 → 0.299 → 0.317 across 2024–26 under one shared vocabulary (the hollow points), — while random pairs stay flat (0.078 → 0.073). Read by section, the rise concentrates where delegation is easiest: — the criticism converges too, but the summary carries most of the headline, and it is the review’s most delegable section. Right: the same pairs read at the unit grain — the overlap between co-reviewers’ sets of charged objects (Plate I’s twelve) also rises after 2023 (0.252 → 0.284, ink). This is the panel where the plate corrects its own first reading, published hours earlier: , and a null that keeps every reviewer’s set size while drawing its contents at random from the year’s mix (the brass dash) rises in step. The excess of choice over size — — is flat. So the convergence is real in the words and unproven in the targets: what co-reviewers say grows alike; what they choose to charge does not measurably follow. That divide has an honest reading the plate cannot prove: expert annotation finds full machine authorship leaving a content fingerprint — AI reviewers’ criticisms overlap with each other seven times more than humans’ (21% vs 3%; Kim et al., arXiv:2605.20668, the “hivemind effect” of Baumann et al., arXiv:2605.03202) — while prose-level assistance would converge the wording and leave the choices alone, which is the pattern this record shows. One neighbouring study reads a shift in ICLR review-text diagnostics at the 2022–23 transition (arXiv:2607.10511); this plate’s own placebo puts the vocabulary’s arrival one cycle later — different instruments, both correlational, both cited rather than reconciled. And one prediction from the LLM-review literature fails here, and is reported as failing: benchmark studies find machine-written reviews lexically flatter (arXiv:2605.25415), but each review’s own lexical diversity sits at a nine-year high in 2026 — type-token ratio 0.638 over a review’s first 200 words, against 0.578–0.598 in every earlier year. The record converges between reviews while growing more varied within them; whatever is homogenizing the juries is not flattening the prose.
Fig. 29d — The structure holds still: three instruments, no narrowing
The mix of criticism spreads slightly, a reviewer’s own apparent spread turns out to be mostly its null, and the distilled reasoning’s co-review gap holds still — nine years, no structural narrowing.
THE CORRECTION’S OTHER HALF — Fig. 29c found the prose converging; these three instruments ask whether the substance followed. It did not: the convergence, so far, lives in the words.
ENTROPY = EVENNESS OF THE MIX, 1.00 = ALL TWELVE EQUAL · MIDDLE, BRASS DASH = N-MATCHED NULL (SAME UNIT COUNTS, OBJECTS DRAWN FROM THE YEAR’S MIX) · GAP = CO-REVIEW MINUS CROSS-PAPER REASONING SIMILARITY · A STRUCTURAL NARROWING WOULD BEND ANY OF THESE DOWN AFTER 2023 — NONE DOES
Fig. 29d
How to read
three panels over the nine-year unit record (952,856 reviewer units, one extraction pipeline throughout), each a way the machine era could have narrowed review’s substance — and did not. Left: the evenness (normalized entropy) of the year’s mix over the twelve objects of scrutiny (ink) and twelve standards (faint), where 1.00 would be a perfectly even docket: . Middle: one reviewer’s own spread — the mean entropy of each reviewer’s object mix (reviewers with five or more units, each normalized by their own ceiling) — rises 0.742 → 0.766. The plate’s first reading said the reviewer “looks at slightly more kinds of things”; an audit the same day held that sentence against the null it deserved, and most of it dissolved: , because reviewers file more units in the LLM years (6.5 → 6.7 per reviewer) and the year’s mix itself grew more even. So this panel affirms nothing about widening — and since its null draws from the year’s mix, it is best read as the left panel’s check at the person grain, not a third independent instrument. What it still rules out is the claim that matters: after 2023 the observed-minus-null excess shows no downward bend. Right: the sub-unit grain — each unit’s distilled reasoning, in embedding space: the gap between co-review pairs of one paper and random cross-paper pairs holds at ≈0.012 in every year, and the year’s overall dispersion is unchanged (0.231 → 0.229). Three owned limits: twelve categories are coarse, so a convergence inside one category is invisible here; the embeddings read the atlas’s distilled reasoning, whose uniform extractor voice is shared by all years — it can mask a converging style, though converging content should still register; and a corpus-level mix says nothing about any single paper. The null pattern is the point: had the machine era narrowed what review looks at, some line here would bend down after 2023.
Fig. 29e — The company the watermark keeps
The marked review is longer, cites less, and hands its paper a slightly higher score than its own co-reviews (firm in 2024–25, marginal in 2026) — and its distinctness is dissolving year by year.
THE AI-LOTTERY DESIGN, WITH A TRANSPARENT INSTRUMENT — Latona et al. found detector-flagged ICLR 2024 reviews scoring higher within the same paper. The paired comparison here replicates that direction for 2024–2026 with the frozen public word list in place of a proprietary detector.
Fig. 29e
How to read
top block: — — small, but on the wrong side of a demanding benchmark: , and marked reviews are the long ones (+58 to +123 words), so length alone predicts the opposite sign. Split by the co-reviews’ own verdict — under a permutation control, because conditioning on the comparison group manufactures a weak-paper gradient out of mean reversion alone — the premium survives in every quality band (+0.07 to +0.28 points over the permuted baseline), running mildly larger for weakly-rated papers in 2024–25, the direction the newest leniency study predicts (Sharma et al., arXiv:2601.20920). Bottom block, with reviews compared only against same-length peers: the marked review cites scholarship less (“et al.” −6.3 points in 2024) and replies less in the rebuttal (−0.08 exchanges) — the profile the ICML study found for machine-modified text — but . The null shape matters as much as the effects: as the watermark fades, the population it marks stops being distinctive — consistent with the residue of Fig. 29b belonging increasingly to ordinary, engaged reviewers who have simply absorbed the vocabulary. Confidence runs mildly higher for marked reviews (+0.03, length-adjusted) in all three years — the one place this record disagrees with the ICML study’s low-confidence profile, and it is reported as the disagreement it is.
CODA
The Archive
the raw reasoning, held against its originals · plate XXX
Plate XXX · The Specimens
Corpus: ICLR 2026 · 232 sampled reviews

Raw reasoning, unretouched

How these were chosen
not at random. For each of the twelve objects of scrutiny, up to twenty reviews were taken that contain at least one unit confidently assigned to that object (label similarity 0.75 or better), deduplicated across objects — 232 reviews in all. The sample therefore covers the taxonomy evenly rather than reflecting how often each object occurs, and it is here to be checked against the originals, not to be counted. Each review is shown as its structured logic flow. Filter by object, standard, or verdict; non-matching units dim. Every title links to the paper's OpenReview page, and each review ID opens the exact review — so you can hold the atlas's reading against the original.
HOW TO READ A SPECIMEN — the italic line is the atlas’s summary of the review; each bordered block below it is one logic unit, the atlas’s structured reading of the reviewer’s prose — not the reviewer’s own words. The four chips name the object of scrutiny, the standard invoked, the verdict, and its provenance (reviewer_explicit = stated outright); the flow then runs inspected → observed → reasoned → judged. Hover any chip for its meaning, and hold the whole reading against the original via REVIEW … OPENREVIEW ↗.
The Finding

No syntax, but a rate card

Three times this atlas searched criticism for a grammar of sequence and logic; three times it found nearly none. What it found instead does not move, and it charges. This page adds nothing new — it only sets the answer in one place, each number quoted from the plate where its caveats live.

Order carries almost nothing
±18%
no argument form’s transition to the next beats chance by more — what little remains is the review form’s template, not a way of thinking. PLATE IV →
script
the arc of a review — warm open, hard middle, hedged close — is the form’s script executed in near-unison, and subject “company” collapses to near-chance under the stricter within-review null. PLATE V →
1.9×
at most, inside the novelty charge itself: the strongest channel from observation to rule; the syllogism’s stages are nearly interchangeable parts. FIG. 7d →
The structure that stands still — and charges
138/144
object–standard couplings sit inside the shuffle floor across nine years: what is inspected is lawfully tied to how it is argued, and the tie does not move. PLATE II →
×2.4
one novelty negative counts more than double the average unit in the tally that reconstructs the score. PLATE VI →
−0.43 pt
the within-paper price of a novelty objection — the same paper, the reviewer who files that charge, that much lower. PLATE XIV →
×0.75
novelty clauses repel each other only within one reviewer (×0.97 across reviewers of the same paper): the one-stamp habit is the writer’s, not the paper’s. PLATE XIII →
39%
of novelty denials ask for nothing at all — every other charge offers a road back 82–92% of the time. PLATE VII →
−10.6 pt
the acceptance gap of papers carrying a novelty charge — the price list runs all the way to the verdict. PLATE XV →

A reviewer’s criticism follows almost no sequence and works only the faintest syllogism. Yet what it inspects is lawfully coupled to how it argues, and what it names has a stable price. Criticism at ICLR is not a grammar. It is a tariff.

Every figure above carries its own caveats where it is drawn · associations, not causes · one venue, one field, nine years · one chosen grain of reading

Appendices
Appendix I · The Lexicon

What every term means

The taxonomy was induced from the data, then named by hand. These are the working definitions behind every label on this page — also available anywhere by hovering the term itself.

Appendix II · Provenance & Method

Where this atlas comes from

Full crosstab — object × reasoning standard
Fig. A1 — The changing words of criticism, 2018–20 vs 2024–26
The vocabulary of criticism moved — partly with the review form itself.

Centroid drift of criticism language is 7–14× the typical year-to-year wobble in categories that are otherwise stable (novelty, empirical scope, clarity) — but read it gently: part of the shift is the review form itself (structured "weakness" fields arrived mid-decade), and the units are one pipeline's paraphrase. The direction that survives the caveat: early-era criticism speaks in verdicts (fails, weak, insufficient), late-era criticism in assessments (limited, concerns, quality, generalizability).

Appendix III · The Null Cabinet

The results that did not survive — kept on purpose

Why this exists
not every analysis in a project this size finds something, and a few found things that a harder look deflated — to near-definitional truths, to artifacts of the review form, or to leans too weak to carry a plate. Deleting them would quietly bias the atlas toward its successes; parading them in the main flow would spend the reader’s trust on weak material. So they live here, each with the claim it once made, the test that deflated it, and where the survivors of that analysis now stand. A null told well is a finding; these are the ones that could not be told well — published anyway.
The archetype mirage, drawn: real roses against forged ones
The forgery blooms: five specialists at ×2.0–3.0 emerge from pure multinomial noise, and .
How to read
each rose has twelve petals, one per reasoning standard; a petal’s length is how far that cluster leans on the standard relative to the whole field (√-compressed), and the dashed ring is ×1 — the rose an interchangeable field would draw. (official reviewers, ≥5 units, 2018–2026). at each reviewer’s own unit volume — profiles that by construction contain no reviewer types at all. Under each rose: its leading standard, that petal’s ratio, and the cluster’s share of profiles; hover any petal for its standard and exact ratio.
Result
the two rows are nearly the same picture — a median profile holds six units, and at that grain k-means carves sampling noise into confident-looking specialists. This is the test that deleted a plate.
Reading it
the mirage is quantified, not total: the real rows sit 1.15× above the multinomial variance floor at the median standard (1.58× at the most dispersed), and the clustering explains 4.8pp more of the real profiles than of the forged ones, a gap that widens with profile length — a faint true signature, far below what the roses’ confidence suggests. Recomputable by scripts/build_mirage_data.py (null seed 46). The cabinet’s entries follow.
Fig. A2 — Acceptance rate by breadth of criticism · kept under glass
Each additional object criticised comes with about five points less acceptance.
RELOCATED FROM THE CONSEQUENCE PLATE — expected by construction, so it lives here rather than spending main-flow attention; the measurement is real and unchanged. No new measurement: take Fig. 15a's twelve with/without splits, sum them per paper into "on how many objects did this paper draw criticism" — how broadly a paper was criticised rather than how harshly — and read the same 13,704 decisions against that count. One number per paper instead of twelve.
Fig. A2
How to read
the horizontal axis is the number of distinct objects criticised, the vertical axis the acceptance rate of the papers at that breadth, and dot size is how many papers stand there.
Result
acceptance falls steeply and then almost linearly — . Few papers escape with fewer than four; . And no decided paper escaped criticism everywhere: breadth zero does not occur, and the two papers criticised on a single object are omitted as too few to chart.
Reading it
association only, and here the two quantities are entangled by construction — a weaker paper gives reviewers more to criticise, so breadth of criticism is partly a measurement of the paper rather than a treatment applied to it.
The archetype mirage — “five kinds of reviewer”
A rose-diagram plate once claimed five reviewer archetypes — Architect, Stress-Tester, Advocate, Gatekeeper, Auditor — from k-means over 98,513 per-reviewer standard profiles, each rose a specialization of ×2.4–3.1 the field. The deflating test is drawn above, at full strength, because it is the cabinet’s centerpiece; the plate was removed on 2026-08-21. What survives is a faint true signature — per-standard variance 1.15× the multinomial floor at the median (1.58× at most), a 4.8pp real-over-noise gap in explained variance that widens with profile length — real, small, and honestly below plate grade. The one live descendant, needing no clustering, is the drift of reasoning-standard shares, now in the Drift plate.
The gauntlet’s slope — breadth of criticism vs acceptance
Acceptance falls ~5 points per additional object criticised, almost linearly (69% at two objects to 27% at nine or more). Expected by construction — a weaker paper gives reviewers more to criticise, so breadth partly measures the paper. Its figure hangs above in this cabinet (relocated from the Consequence plate, 2026-08-25) for its residual facts: the near-linearity itself, that no decided paper escaped criticism everywhere, and that the modal paper is criticised on seven of twelve objects.
Whose rating the chair echoes — a lean, not a rule
In the Borrowed Verdict, the panel’s uniquely-lowest rater is echoed most in 16% of panels against 13% for the uniquely-highest — and in half of panels the echoed rating ties an extreme under the coarse scale, so no verdict is possible at all. The direction claim that plate makes stands on its 60/40 rejection tilt instead; this per-reviewer lean alone would not have carried it.
No syntax of argument forms
Plate IV’s finding in full: once a review’s own mix is fixed, no argument form predicts the next beyond ±20% — judgment, as written, has almost no grammar. Kept in the main flow because the absence is informative; indexed here because it is the project’s cleanest true null. Addressee: work that assumes review-internal argument structure.
Asked versus asserted — the interrogative null
Whether an objection is phrased as a question or an assertion makes no detectable difference to its fate after the rebuttal (the Fate of an Objection’s second figure). Reported in place as a null; indexed here.
The advocate’s softness — near-definitional
An early “advocate reviewers soften more” reading collapsed on inspection: the defining standard of the advocate profile is the praise standard, so the claim was close to circular and was never published as a finding.
Reading it
the common thread: three deflators account for every entry — the review form’s own structure, volume-plus-noise nulls, and definitional circularity. They are the same three checks every surviving plate had to pass; this cabinet is what passing looks like from the other side.
The Door

The sentence you were handed

Most readers arrived holding one line from a review — "limited novelty," "insufficient baselines." By now that line has a docket. Pick it here and the page returns to the plate that takes it apart: what reviewers had actually observed when they wrote it, the law they were applying, and the remedy they most often prescribe. The three links under the grid go to what reviewers ask for, what each criticism costs in score, and what happened to authors who answered.

The whole corpus, printed on one sheet
1,420,178 acts of judgment
every mark on this sheet is one real unit of reviewer reasoning — the same units the plates have been reading, gathered back into a single impression · nine years of ICLR