“Limited novelty: Core idea — early-commitment to promising partial solutions — closely resembles existing beam/pruning and adaptive-consistency techniques; incremental contribution.”
Four parts, one unit. This exact dissection was performed 410,586 times on ICLR 2026, and a million more across nine years of discussion — every figure below is a count of these. The note that follows says what a machine’s reading costs, why this particular cut is a choice, and how to check both.
Open-weight language models read each review and split it into the separate judgments it makes. One unit records four things: the part of the paper inspected, what was observed there, the standard the reviewer judged it by, and the verdict reached. A typical review yields five or six. Every figure on this page is a count of units, so the labels are a machine's reading of the text — consistent across a million reviews, but not ground truth. The grain is a choice, too: these four parts, and the 12 × 12 taxonomy the plates sort them into, are one induced way of carving reasoning, held fixed throughout — internally consistent by the Method page's checks, but not a carving whose aptness any plate can test. A different scheme would draw a different map. Units link back to the review they came from; check them.
The reviews were read twice, in two different ways, and the plates use whichever fits the question. ICLR 2026 was read one review at a time (410,586 units), so each unit can point at the line of the review it came from. 2018–2026 was read a second way — each paper's whole discussion in a single pass, including the replies written after the author's rebuttal and the area chair's meta-review (1,009,592 units) — which is what makes it possible to see when a judgment was made and whether it later changed. Neither is a subset of the other. Each plate's header names the years it uses; the Method page gives both pipelines in full. One more thing worth knowing before the sequence starts: the order of the plates is an editorial arrangement, made after the analyses existed, not the order in which the work was done — the Method page carries a plate-by-plate ledger of which questions were written down before the numbers existed, which plates are frankly exploratory, which we simply have no record for, and one claim that was published wrong and corrected in place. Finally, the stage-setting plates travel as overtures at the heads of the acts they serve: the calendar of one full review season opens Act V, one warning about review timing — reviews filed weeks early run harsher and longer than deadline-day ones, and no plate controls for filing time — opens Act VII as the Tide, and the venue's own census opens Act VIII.
in 2026 alone
Object → standard → verdict
DERIVATION 12 × 12 pairings − the merit-recognition column (12) − 1 cell under 40 units = 131; each has a negative share above 50%
VERIFIED D10 number audit, 2026-08-24 outside the merit-recognition column condemn more often than not; inside that one column, criticism falls to 14–35% for eleven of the twelve objects.
DERIVATION mean presentation sub-score of reviews with ≥1 negative clarity unit minus reviews without, computed within each overall-rating level and weighted-averaged: −0.3465 on the 1–4 scale
VERIFIED checked against shipped JSON, 2026-08-27 — and the one place they part is itself a finding.
The law is often borrowed
DERIVATION count of the 12×12=144 (object, standard) cells where |delta| ≤ floor95 (0.0405) = 138
VERIFIED checked against shipped JSON, 2026-08-25 — and what little moved was mostly ground the novelty standard gave up.
Each standard has a voice
One question, three windows
The next two plates ask one question through three windows — when a review is read in the order it was written, does position carry information? — and a reader meeting the three analyses in a row may fairly ask what separates them. The answer is the window. Below is the same review three times: each stack is one review's units in written order, coloured by the section of the review form they sit in — strengths, weaknesses, questions, the same colours the plates ahead use for positive, negative, and hedged verdicts. The bright part is what that plate reads; the fine print beneath names the chance level its ×-numbers are measured against. Click a panel to jump to its plate. And one adjustment of expectations, made honestly up front: these three windows mostly find weak or negative results. The strongest order in a review turns out to belong to the review form, not the reviewer; the grammar hunt comes back nearly empty; what survives — who gets the first word of the criticism — is real but modest. The plates keep these results because the absences are informative: knowing that judgment has almost no syntax is worth as much as a syntax would have been.
The search for a syntax finds almost none
Plate III sorted every unit into six argument forms. Do the forms care what comes next? For every review, count which form follows which between consecutive units, and compare against a null that preserves the review's own composition — so a common form cannot masquerade as an attractive one.
Criticism is written in an order
DERIVATION share of reviews mentioning novelty that mention it exactly once = 0.9655 ≈ 97%
VERIFIED checked against shipped JSON, 2026-08-25.
The score counts the verdicts — it barely reads the topics
What a charge actually says
Act II watched one reviewer write — the order, the arc, the early verdict. This act stops watching the writer and cross-examines the writing. Tell a first-year researcher their paper "lacks novelty" — or rigor, or clarity — and they learn nothing. This plate takes such a sentence apart. Because each unit stores the reviewer's observation and their reasoning separately, every negative unit can be split into three: the ground — what the reviewer actually saw in the paper; the rule — the general principle they applied to it, quoted in their own words; and the remedy — what they asked the author to do, if anything. (The courtroom words are used consistently from here on: a negative unit is a denial or a charge, and the principle behind it is a rule or a law.) One more borrowed word: a docket is all the criticism aimed at one of the twelve objects of scrutiny from Plate I, so "the novelty docket" means every criticism of novelty. Pick one below and it opens as ground → rule → remedy. A single criticism can invoke several rules at once, and nothing here forces it into one.
DERIVATION min across dockets = clarity 26.4%, max = novelty 40.4%
VERIFIED checked against shipped JSON, 2026-08-25.
DERIVATION lift = observed transition count ÷ (row marginal × column marginal / n); merged-grain maximum 1.88× (“cannot tell what is new → rigor doubts imported”); raw-grain χ² = 3,690 (ground→warrant) and 4,601 (warrant→ask incl. no-ask), n = 30,447, Cramér’s V 0.116 / 0.137
VERIFIED recomputed from shipped JSON, 2026-08-27. The stages of the argument are nearly interchangeable parts.
What “mere combination” actually says
Plate VII found the combination rule — assembling known components is not novel unless the integration itself is. This plate opens the clause: what, concretely, is the unless? And what is being called a combination in the first place? Sub-clustering the rule’s own sentences answers the first; the units’ observations answer the second.
DERIVATION sum of .share over the six non-dead-end exceptions (excludes "The threshold, restated" and "The reproduction floor") = 0.7595 ≈ 76%
VERIFIED checked against shipped JSON, 2026-08-25 — added value, new insight, differentiation, a new mechanism, an isolated contribution, innovation beyond the integration — and two clauses that offer none.
The judgment that needs no referent
Plate VIII opened one clause of the novelty rule from the inside — what redeems “mere combination.” This plate opens a different clause the same way: every novelty objection implies a comparison — this work against the prior it allegedly repeats. We went looking for the canon: the named prior works that kill novelty. The search came back nearly empty, and the emptiness describes the standard itself. A judgment that cites nothing is not automatically careless — novelty can be judged from the reviewer’s internal map of the field, which is much of what expertise is — but a judgment made that way cannot be checked from the text, and its burden of proof moves to the author. What follows measures that mode of judgment: how common it is, and which way it has moved as the venue grew.
DERIVATION named-citation share fell from 20.15% (2018) to 10.77% (2026), roughly half
VERIFIED checked against shipped JSON, 2026-08-25, under both instruments tried.
Criticism has a common coin
Take any criticism and search every other paper’s reviews for the most similar one — its nearest twin. Some criticisms have no close twin: they could only have been written about this paper. Others recur almost word for word across many papers — the field’s standard formulas. Finding formulas is not an accusation: a shared formula is how a discipline applies the same standard to everyone ("release the code" has to read the same everywhere to be a standard at all). The question is which criticisms are formulas and which are written for the paper at hand — and Fig. 10b’s answer is that the two kinds carry very different weight. After three plates of nine-year corpora, the window narrows again: the embedding search reads ICLR 2026 alone — 293,671 negative units.
DERIVATION the median nearest-other-paper-twin cosine similarity across all embedded 2026 negative units is 0.8114, which rounds to 0.81
VERIFIED checked against shipped JSON, 2026-08-25.
What reviewers actually ask for
When a review is split into units, each criticism keeps whatever concrete suggestion the reviewer wrote next to it — the same fix field Plate VII read as remedies. 85% of criticism — and three quarters of everything reviewers set down — arrives with such a repair attached. This plate asks what those repairs actually say: 80,000 of them are embedded and clustered into a working taxonomy of the discipline's demands. Read at the family grain of Plate VII, one kind dominates — over two thirds of everything requested is work of research: new evidence in some form.
DERIVATION the top-ranked of 16 merged repair-type clusters ("Extend the evidence") holds 22.15% of the 80,000 sampled 2026 suggested-fixes, clearly ahead of the next two clusters (8.59%, 8.48%)
VERIFIED checked against shipped JSON, 2026-08-25 — and its whole family, the work of research, is over two thirds of everything asked.
DERIVATION “extend the evidence” leads at every k, its share moving 22.2–27.0% and always ≥1.8× the runner-up; the work-of-research family sum moves only 66.8–71.3% across the same 4× range of k; k=22 reproduces the shipped partition exactly (ARI 1.000, seed 7)
VERIFIED k-sweep run 2026-08-27, adversarial-audit cycle), so the stabler finding is . And one absence, drawn as the empty dashed shelf at the bottom: , the same absence Fig. 7c found at coarse grain. On first view the figure assembles as sediment: each falling grain is one sampled ask, settling into its shelf — and nothing ever lands on the redesign shelf.
Where scrutiny turns hostile — and which criticism arrives with a way out
DERIVATION summing all 410,586 2026 units by object: related_work is 9,033/10,306 = 87.65% negative, novelty is 9,409/11,440 = 82.25% negative
VERIFIED checked against shipped JSON, 2026-08-25 — is judged almost only in the negative.
Which laws are filed together
Plate VII collected the rules reviewers state in their own words — proofs must be rigorous, notation must be defined, novelty requires differentiation from prior art. A single review usually states several. This plate asks: which rules tend to appear in the same review together — and what that pattern says about the person writing. The logic is simple: if each rule appeared only because of the paper’s own faults, any two rules would share a review at plain chance rate, so every departure from chance points at something else. Departures are measured with the ×chance lift used throughout the atlas — 2× means a reviewer who states one rule states the other twice as often as chance — across all 191,946 initial reviews. Two kinds of departure appear. Rules about rigour and presentation cluster in the same review, for two reasons at once: a paper that earns one such complaint tends to earn the others, and a careful reader is careful about everything (the two are told apart by checking the other reviewer of the same paper — the caption walks through it). The novelty complaint does the opposite: a reviewer who has written “not novel” once does not write it again under another heading, while the other reviewer of the same paper is just as likely to write their own. The clustering is partly the paper’s doing; the once-only habit belongs to the reviewer alone. In the figures below the rules are called laws and a review’s set of them its charge sheet — the atlas’s courtroom image, and nothing more.
DERIVATION across all readable rule-pairs the median lift is ×1.185 (≈1.19); the strongest attraction is clarity's precision rule × theory's rigor rule at ×3.64; the strongest repulsion is method design's novelty rule × novelty's differentiation rule at ×0.658
VERIFIED checked against shipped JSON, 2026-08-25 — each with its 95% interval and its cross-reviewer twin.
What a fault costs, and whose law decides
A review ends in one number. This plate asks three questions that number hides. First, what does each objection cost — how much lower does a reviewer score a paper when they raise it? Comparing reviews of different papers would answer nothing: weak papers attract both criticism and low scores. So the cost is measured between reviewers of the same paper, also holding fixed how much criticism each review contains overall — what survives is the price of that particular objection, the plate’s tariff. Second, is a fault’s price tied to whether it can be repaired — do the costly objections arrive with a prescribed fix, or without one? Third, when two reviewers examine the same object of the same paper, does their agreement depend on the reasoning standard each brought to it — the laws of Plate VII? And fourth, one level finer and nine years wide: when a criticism states its rule in words, does the stated rule set the price — or does the same rule cost differently depending on where it is aimed?
DERIVATION novelty's within-paper coefficient is −0.429; the next-largest negative-signed coefficient among objects whose 95% interval excludes zero is problem framing at −0.179 (robustness/compute-cost run positive-signed, a different kind of association, so excluded); 0.429 ÷ 0.179 = 2.40×
VERIFIED checked against shipped JSON, 2026-08-25 — while . Robustness and compute-cost objections carry a positive sign: they are the objections of reviewers who found nothing worse to say. Read these as associations, not effects. Nobody assigned the objections at random: a reviewer who is already unconvinced is the one who raises novelty, so the number measures which objection travels with a low score, not what raising it would do. ⚖ reviews are not interchangeable draws — early-filed ones run harsher and longer, uncontrolled throughout — the tide · overture to act VII
Fig. 14b — Each object placed by the share of its negative units that arrive with a prescribed repair (horizontal) against its within-paper tariff (vertical). Circle area tracks how many criticisms the object drew, fill repeats Fig. 14a's color — red = costs rating points, green = travels with higher ones, gray = interval crosses zero — and each label is inked in its circle's color, with a thin line tying the two — novelty alone in white, the plate's finding. The bill collects in one corner: the one objection that rarely comes with a fix — — is also by far the costliest. Cheap objections are prescriptions; the expensive one is a sentence. The unlabeled points crowding the near-zero center — — are the objections that neither cost nor spare; hover any point for its numbers.
Which criticism travels with rejection
DERIVATION novelty.gap = −0.1058 × 100 = −10.6 points (accept-rate of papers with a novelty negative unit minus papers without one)
VERIFIED checked against shipped JSON, 2026-08-25; compute-cost and robustness criticism travel with nothing.
DERIVATION 5-fold stratified CV, logistic regression, 39,484 decided papers 2018–2026 (2019 carries no decision records in this corpus); tally = per-object counts of pre-rebuttal negative units; text = mean-pooled bge-small embedding of the same units’ judgments; rating_ref = the panel’s scores, shown as reference
VERIFIED checked against shipped JSON, 2026-08-27 — and the tally’s share has been falling for nine years.
One review season, replayed
Before asking whether judgment moves when people meet, watch the calendar the meetings happen on. Every public act of the ICLR 2026 cycle, on its real calendar day: 19,814 submissions crest at the deadline, 75,859 reviews land in a two-day burst, 180,146 discussion comments answer them, 5,216 papers withdraw — most in the week after the reviews arrive — and on a single day in January, 14,175 decisions fall at once. This overture replays a single season; the plates that follow read nine years of such seasons.
Every conversation has a skeleton
DERIVATION returned (20,453) ÷ n (46,748) = 43.8% of 2025 review threads where the reviewer wrote again
VERIFIED checked against shipped JSON, 2026-08-25, then collapsed in 2026 — the year submissions grew 70% and the number of reviews grew 62%.
One event, counted three ways
The next three plates watch a single event — after the author’s rebuttal, did the reviewer’s written judgment soften? — and differ only in what they divide it by. The Rebuttal counts per unit written after the response: of what got written again, what moved. The Moves counts per author–reviewer exchange, joined to what the author did: which moves travel with movement. The Fate of an Objection counts per objection originally raised — the only denominator that can see the objections never spoken of again, and they turn out to be the large majority. Same event, three denominators; the denominator decides what each number can mean.
sees only what was written again
sees what the author did
the only window that can see the silence
Author responses barely move the written judgment
130,650 units sit after the author response (2018–2026) — written by the reviewers who spoke again at all; most never do (Plate XVI). When a written judgment moves, it moves toward reinforcement 2.4 times more often than concession — and every kind of response, fresh experiments included, reinforces at least twice as often as it softens. Two limits, stated up front: this plate reads what reviewers wrote, so a score quietly raised without a word is invisible — the corpus keeps no rating history — and every recorded fate is a model's reading of the text.
DERIVATION 0.0281 ÷ 0.0116 = 2.4× (share of low-scorer post-rebuttal units that soften/reverse, vs. the high-scorer's share, across 9,329 split panels)
VERIFIED checked against shipped JSON, 2026-08-25.
What authors do, and what actually moves
Every author reply in a reviewer’s thread, scanned for seven quotable moves — the concession, the contest, the delivered experiment, the promissory note — and joined to whether that reviewer’s post-response units record any weakened or reversed judgment. An exchange exists only where the reviewer wrote again at all — most never do (Plate XVI). Across all 67,290 exchanges, SOURCE moves-data.json → soft_base, n_pairs
DERIVATION soft_base = 0.0597 = 5.97% ≈ 6% of the 67,290 author-reply/reviewer-response exchange pairs where the reviewer's post-response judgment softened or reversed
VERIFIED checked against shipped JSON, 2026-08-25 — that is the number every move below has to beat.
DERIVATION sum of .raised over all 12 objects in LIFECYCLE.objects = 551,595
VERIFIED checked against shipped JSON, 2026-08-25
Most objections are never spoken of again
Follow every objection — a reviewer's initial negative unit on one object — through the discussion phase: did the same reviewer return to that object after the authors responded, and if so, did the judgment soften or harden? The dominant fate, for every kind of objection, is silence.
[questions] field), grey for objections asserted (in [weaknesses]), violet for objections the reviewer wrote in both fields. Every row runs on its own 0→max scale. If the question mark changed an objection's fate, the brass and grey dots would sit apart.Dissent, and the kinder judge
Reviewers of the same paper, read as a panel — four questions in one plate. Where do co-reviewers take opposite stances (Fig 20a)? How does the meta-reviewer’s tone differ from the panel they summarize (the three numbers beside it)? When the final decision disagrees with the panel’s lean, which way does it err (Fig 20b)? And how much of the ground does a panel actually walk (Fig 20c)?
DERIVATION 919 papers accepted despite a below-line panel mean ÷ 338 rejected despite an above-line mean = 2.72, rounds to 2.7×
VERIFIED checked against shipped JSON, 2026-08-25.
The overruled — public exemplars
Watch a panel think
First, how the 17,848 panels of ICLR 2026 usually end — then fourteen real deliberations, written out like music. Each reviewer is a staff and each unit a note, placed in the order it was reasoned; the dashed barline is where the authors respond, and what follows it carries the fate of each judgment — hardened, softened, or reversed. Faint verticals bind reviewers who touched the same object: ink where they agree, sienna where the same object drew opposite verdicts. Hover any note to read the unit.
DERIVATION 15,289 panels with no recorded strengthen/weaken/reverse ÷ 17,848 rated panels × 100 = 85.66%, rounds to 85.7%
VERIFIED checked against shipped JSON, 2026-08-25 — and when judgments do move after the response, .
Everything else in this atlas averages thousands of panels; this figure does the opposite. Below are individual papers — fourteen real 2026 discussions, hand-picked as extreme, legible specimens of the five endings above (not a random sample; the pull-quotes on each card are the panel's final ratings). Open a case to read its full score.
The judge above the judges
Above every panel sits an area chair who writes a meta-review — and those were read the same way: SOURCE panel-data.json → meta.meta_reviewer.n
DERIVATION raw count of extracted meta-review logic units, n = 56,736
VERIFIED checked against shipped JSON, 2026-08-25. The higher court has its own grammar: it does not re-litigate the evidence, it weighs the verdicts. And at the margin, where panels sit just below the accept line, its choices reveal which objections it declines to forgive.
Fig. 22c — Share of all meta-review units on those same borderline papers that praise each object, for lifted papers against rejected ones. On a lifted paper roughly three-quarters of the meta-review is praise, and its favorite subjects are the numbers and the positioning — "the improvements are significant, the comparisons honest" — while on a rejected borderline paper praise of any kind nearly vanishes. Read as the language of justification, not independent evidence: the meta-review is written by the person who has already decided.
Whose words reach the decision
Match every meta-review unit to its nearest reviewer unit on the same paper — nearest in wording, not in time — across the 8,439 papers of 2018–2026 with a meta-review of three or more units and at least two official reviews. Three tiers fall out: near-copies, echoes, and the area chair’s own words — and the borrowing has a direction.
DERIVATION own = 0.7306 → 73%; copy = 0.0047 → 0.5%, share of meta-review units by cosine-similarity tier to their nearest reviewer unit
VERIFIED checked against shipped JSON, 2026-08-25. But criticism is borrowed more than praise — at every cutoff.
The early review is a different document
What each score sounds like
How repeatable is a score?
Treat a paper's rating as a measurement and ask what any instrument must answer: if the measurement were repeated — same paper, another qualified reviewer — how much would it move? The NeurIPS consistency experiments answered by having two independent committees review the same papers; this is the observational shadow of that experiment, computed from every co-review of the same paper across nine years of ICLR.
DERIVATION the intraclass correlation for 2026 is stored directly as 0.18
VERIFIED checked against shipped JSON, 2026-08-25).
Fig. 25e — Every square is one of the 5,359 papers accepted at ICLR 2026, ordered by panel mean from the strongest accept down. Press the button and the model deals one alternate conference: brass squares survive the redraw, dark ones do not — and their seats would largely be taken by papers rejected this time (a quarter of rejections flip the other way). Each press is a new deal; no two alternate conferences are the same, which is the point.
ICLR itself, 2018 → 2026
Before asking what belongs to an era or a territory, start with the venue itself, apart from any of the review analysis. These are plain counts from ICLR's public record on OpenReview: how many papers were submitted, how many were accepted, how long the reviews ran. , and that growth is the backdrop for everything in this atlas.
Full census table
Nine years of shifting scrutiny
DERIVATION (0.0831 − 0.1319) / 0.1319 = −37.0% relative drop in clarity's share of review units
VERIFIED checked against shipped JSON, 2026-08-25 and compute cost's rose 67%. What the series cannot separate is reviewers changing from submissions changing: it measures what got written, not why. And the field is hardening: the share of post-response judgments that weaken or reverse has halved, 8.2% in 2018 to 4.2% in 2026 — though 2018 rests on only 1,239 post-response units.
Novelty's stated definitions look the same in 2018 as in 2026
The Drift plate showed enforcement changing — which objections are raised, how often, at what price. This plate asks about the statute itself. It was run expecting to find a change: nine years spanning the arrival of large language models seemed likely to have rewritten what counts as a new contribution. The result is negative, and is reported as a negative result. The test: fit the code of novelty on 2026 sentences alone, then hold every earlier year against it. If the law had been rewritten, 2018 should sit measurably farther from the 2026 code than 2026 does from itself.
DERIVATION 2018's mean embedding distance to the 2026 code (0.159) compared to the code's own held-out-half calibration distance (0.1516 → 0.152)
VERIFIED checked against shipped JSON, 2026-08-25.
Choose a field. Face its tribunal.
DERIVATION dev[theory] = 0.2541 → +25.4pp above the venue baseline; dev[baselines_ablations] = −0.3051 → −30.5pp, rounding to −31pp
VERIFIED checked against shipped JSON, 2026-08-25.
A watermark appears in 2024 — and fades while the practice spreads
DERIVATION the signature’s share falls inside every fixed word-count band from 2024 to 2026 (500–699 words: 10.4% → 4.4%; even 700+: 10.2% → 7.1%) while the mean review lengthened, 406 → 425 words
VERIFIED adversarial-audit cycle 2026-08-27, builder-reproduced same day. And the newest outside estimate points the other way: a January 2026 study applying the ICML likelihood method reads 26.7% of ICLR 2025 reviews as LLM-involved (Sharma et al., arXiv:2601.20920), rising while this tracer’s mark fell. The one year the record can calibrate against says the same: in 2024, when an independent method estimated one review sentence in ten machine-modified, this tracer marked only 7.8% of reviews — the watermark understates the practice even at its darkest.
DERIVATION 0.07785 − 0.01644 = +0.0614, i.e. +6.1 points of excess over the 2018–2022 linear trend
VERIFIED checked against shipped JSON and recounted via direct SQL, 2026-08-27, . . A negative-control list of ten ordinary reviewer words (interesting, unclear, convincing…), run through the identical machinery, shows no spike — it drifts mildly below its trend as reviews shortened, which makes the marker spike conservative, not inflated. The excess is a floor, not a count: a review whose machine-assisted text happens to avoid all thirteen words is invisible here.
DERIVATION document frequency in 2026 divided by the word’s own 2018–2022 linear extrapolation; delves: 0.00037 / 0.00104 = 0.35
VERIFIED checked against shipped JSON; SQL recount finds delve-family reviews falling 551 → 138 from 2024 to 2026 while the corpus grew 2.7×, 2026-08-27 — undershooting it, as if the word had been scrubbed — along with showcases (×0.51) and showcasing (×0.98), while underscoring (×253, from near-zero), underscores (×6.5), meticulous (×5.1) and commendable (×4.7) persist. One list, two fates: the vocabulary of the 2023-era models burned bright and burned out; a quieter residue is still spreading. For scale, the biggest 2026 vocabulary shifts of all are the field’s own subject matter — qwen appears in 9.2% of 2026 reviews, llama in 6.4%, deepseek in 2.6%, against essentially zero before — which is why the tracer is restricted to style words: topic drift is real, loud, and not evidence of who held the pen.
DERIVATION the identical fixed-vocabulary pipeline run per section: summary-only within-paper median 0.2801 → 0.3102 → 0.3357 (+20%); weaknesses+questions-only 0.1608 → 0.1720 → 0.1708 (+6%, then flat); mean section lengths 89/87/93 words (summary) vs 273/294/275 (criticism) — near-constant within each section, so the split is not a length artifact
VERIFIED reproduced by build_llmtrace_mix.py, 2026-08-27 — the criticism converges too, but the summary carries most of the headline, and it is the review’s most delegable section. Right: the same pairs read at the unit grain — the overlap between co-reviewers’ sets of charged objects (Plate I’s twelve) also rises after 2023 (0.252 → 0.284, ink). This is the panel where the plate corrects its own first reading, published hours earlier: , and a null that keeps every reviewer’s set size while drawing its contents at random from the year’s mix (the brass dash) rises in step. The excess of choice over size — SOURCE llmtrace-data.json → mix.jaccard_size_null[year].excess
DERIVATION observed median co-reviewer Jaccard minus a 20-simulation null holding each reviewer’s set size and drawing contents from the year’s object mix; 2026: 0.2837 − 0.2418 = +0.042
VERIFIED checked against shipped JSON, 2026-08-27 — is flat. So the convergence is real in the words and unproven in the targets: what co-reviewers say grows alike; what they choose to charge does not measurably follow. That divide has an honest reading the plate cannot prove: expert annotation finds full machine authorship leaving a content fingerprint — AI reviewers’ criticisms overlap with each other seven times more than humans’ (21% vs 3%; Kim et al., arXiv:2605.20668, the “hivemind effect” of Baumann et al., arXiv:2605.03202) — while prose-level assistance would converge the wording and leave the choices alone, which is the pattern this record shows. One neighbouring study reads a shift in ICLR review-text diagnostics at the 2022–23 transition (arXiv:2607.10511); this plate’s own placebo puts the vocabulary’s arrival one cycle later — different instruments, both correlational, both cited rather than reconciled. And one prediction from the LLM-review literature fails here, and is reported as failing: benchmark studies find machine-written reviews lexically flatter (arXiv:2605.25415), but each review’s own lexical diversity sits at a nine-year high in 2026 — type-token ratio 0.638 over a review’s first 200 words, against 0.578–0.598 in every earlier year. The record converges between reviews while growing more varied within them; whatever is homogenizing the juries is not flattening the prose.
DERIVATION Shannon entropy of the year’s 12-object unit counts ÷ log₂12; 2018: 0.9155 → 2026: 0.9395 (HHI 0.119 → 0.109)
VERIFIED checked against shipped JSON, 2026-08-27. Middle: one reviewer’s own spread — the mean entropy of each reviewer’s object mix (reviewers with five or more units, each normalized by their own ceiling) — rises 0.742 → 0.766. The plate’s first reading said the reviewer “looks at slightly more kinds of things”; an audit the same day held that sentence against the null it deserved, and most of it dissolved: , because reviewers file more units in the LLM years (6.5 → 6.7 per reviewer) and the year’s mix itself grew more even. SOURCE llmtrace-data.json → mix.attention_null[year]
DERIVATION observed mean entropy minus a 5-simulation null holding each ≥5-unit reviewer’s n and drawing objects iid from the year’s object mix; excess −0.046 (2018) → −0.038 (2026), range −0.053 to −0.038 with no monotone trend; after 2023: −0.044 → −0.048 → −0.047 → −0.038
VERIFIED reproduced by build_llmtrace_mix.py, 2026-08-27 So this panel affirms nothing about widening — and since its null draws from the year’s mix, it is best read as the left panel’s check at the person grain, not a third independent instrument. What it still rules out is the claim that matters: after 2023 the observed-minus-null excess shows no downward bend. Right: the sub-unit grain — each unit’s distilled reasoning, in embedding space: the gap between co-review pairs of one paper and random cross-paper pairs holds at ≈0.012 in every year, and the year’s overall dispersion is unchanged (0.231 → 0.229). Three owned limits: twelve categories are coarse, so a convergence inside one category is invisible here; the embeddings read the atlas’s distilled reasoning, whose uniform extractor voice is shared by all years — it can mask a converging style, though converging content should still register; and a corpus-level mix says nothing about any single paper. The null pattern is the point: had the machine era narrowed what review looks at, some line here would bend down after 2023.
DERIVATION papers holding both a marked and an unmarked review; mean(marked ratings) − mean(unmarked ratings), averaged over papers (1,919 / 2,563 / 2,189 papers)
VERIFIED checked against shipped JSON, 2026-08-27 — small, but on the wrong side of a demanding benchmark: , and marked reviews are the long ones (+58 to +123 words), so length alone predicts the opposite sign. Split by the co-reviews’ own verdict — under a permutation control, because conditioning on the comparison group manufactures a weak-paper gradient out of mean reversion alone — the premium survives in every quality band (+0.07 to +0.28 points over the permuted baseline), running mildly larger for weakly-rated papers in 2024–25, the direction the newest leniency study predicts (Sharma et al., arXiv:2601.20920). Bottom block, with reviews compared only against same-length peers: the marked review cites scholarship less (“et al.” −6.3 points in 2024) and replies less in the rebuttal (−0.08 exchanges) — the profile the ICML study found for machine-modified text — but . The null shape matters as much as the effects: as the watermark fades, the population it marks stops being distinctive — consistent with the residue of Fig. 29b belonging increasingly to ordinary, engaged reviewers who have simply absorbed the vocabulary. Confidence runs mildly higher for marked reviews (+0.03, length-adjusted) in all three years — the one place this record disagrees with the ICML study’s low-confidence profile, and it is reported as the disagreement it is.
Raw reasoning, unretouched
No syntax, but a rate card
Three times this atlas searched criticism for a grammar of sequence and logic; three times it found nearly none. What it found instead does not move, and it charges. This page adds nothing new — it only sets the answer in one place, each number quoted from the plate where its caveats live.
A reviewer’s criticism follows almost no sequence and works only the faintest syllogism. Yet what it inspects is lawfully coupled to how it argues, and what it names has a stable price. Criticism at ICLR is not a grammar. It is a tariff.
Every figure above carries its own caveats where it is drawn · associations, not causes · one venue, one field, nine years · one chosen grain of reading
What every term means
The taxonomy was induced from the data, then named by hand. These are the working definitions behind every label on this page — also available anywhere by hovering the term itself.
Where this atlas comes from
Full crosstab — object × reasoning standard
Centroid drift of criticism language is 7–14× the typical year-to-year wobble in categories that are otherwise stable (novelty, empirical scope, clarity) — but read it gently: part of the shift is the review form itself (structured "weakness" fields arrived mid-decade), and the units are one pipeline's paraphrase. The direction that survives the caveat: early-era criticism speaks in verdicts (fails, weak, insufficient), late-era criticism in assessments (limited, concerns, quality, generalizability).
The results that did not survive — kept on purpose
DERIVATION share of between-profile variance the 5-cluster k-means partition explains: 0.3578 (35.8%) for the 98,513 real reviewer profiles, 0.3102 (31.0%) for the synthetic multinomial-noise profiles
VERIFIED checked against shipped JSON, 2026-08-25.
- The archetype mirage — “five kinds of reviewer”
- A rose-diagram plate once claimed five reviewer archetypes — Architect, Stress-Tester, Advocate, Gatekeeper, Auditor — from k-means over 98,513 per-reviewer standard profiles, each rose a specialization of ×2.4–3.1 the field. The deflating test is drawn above, at full strength, because it is the cabinet’s centerpiece; the plate was removed on 2026-08-21. What survives is a faint true signature — per-standard variance 1.15× the multinomial floor at the median (1.58× at most), a 4.8pp real-over-noise gap in explained variance that widens with profile length — real, small, and honestly below plate grade. The one live descendant, needing no clustering, is the drift of reasoning-standard shares, now in the Drift plate.
- The gauntlet’s slope — breadth of criticism vs acceptance
- Acceptance falls ~5 points per additional object criticised, almost linearly (69% at two objects to 27% at nine or more). Expected by construction — a weaker paper gives reviewers more to criticise, so breadth partly measures the paper. Its figure hangs above in this cabinet (relocated from the Consequence plate, 2026-08-25) for its residual facts: the near-linearity itself, that no decided paper escaped criticism everywhere, and that the modal paper is criticised on seven of twelve objects.
- Whose rating the chair echoes — a lean, not a rule
- In the Borrowed Verdict, the panel’s uniquely-lowest rater is echoed most in 16% of panels against 13% for the uniquely-highest — and in half of panels the echoed rating ties an extreme under the coarse scale, so no verdict is possible at all. The direction claim that plate makes stands on its 60/40 rejection tilt instead; this per-reviewer lean alone would not have carried it.
- No syntax of argument forms
- Plate IV’s finding in full: once a review’s own mix is fixed, no argument form predicts the next beyond ±20% — judgment, as written, has almost no grammar. Kept in the main flow because the absence is informative; indexed here because it is the project’s cleanest true null. Addressee: work that assumes review-internal argument structure.
- Asked versus asserted — the interrogative null
- Whether an objection is phrased as a question or an assertion makes no detectable difference to its fate after the rebuttal (the Fate of an Objection’s second figure). Reported in place as a null; indexed here.
- The advocate’s softness — near-definitional
- An early “advocate reviewers soften more” reading collapsed on inspection: the defining standard of the advocate profile is the praise standard, so the claim was close to circular and was never published as a finding.