Policy
- Send the exact row ID, the disputed field, and a stable primary-source locator through a repository issue ↗ or email.
- Every report is logged publicly the same calendar day it is received: either as a verified correction or as an explicit under-review entry. Silence is not a resolution.
- A verified correction updates the source row, all mechanically derived counts, generated pages, and the revision history in one reviewable change. The report is credited in the affected row where consent permits.
- If evidence remains ambiguous, the row is weakened to an explicit ambiguity rather than retained by confidence or rewritten criteria.
Public revision history
The full record is also embedded beside the census rows; this route exists so a citation, post, or correction request always has one stable destination.
- 2026-08-27 — Census established. Criteria v1 inclusion wording was committed and is now locked in repository history; that lock is a reproducibility record, not an independent preregistration. The four starting cases entered as under_review pending primary-source examination.
- 2026-08-27 — First examination pass completed: 19 rows examined (one over the ~18 budget estimate, recorded here), 9 candidates excluded by rule, 15 surfaced candidates left unexamined. Pre-publication schema clarifications, made before any count was public: the ABSENT definition no longer presupposes comparability (that axis is counted by M); joint_scope value printed_pairwise_only generalized to printed_partial_stack. Corrections to the starting brief: the single BELLS row split into three artifacts (2024 framework — excluded as a framework paper; 2025 misuse-detection evaluation; 2026 BELLS-O), and the artifact remembered as "MSBench with two pairwise ensemble rows" is actually UnsafeBench (arXiv 2405.03486), which prints six OR-ensemble rows but evaluates image-platform safety classifiers and therefore fails frozen domain criterion 5 — recorded prominently under exclusions. LlamaFirewall added from the snowball hop.
- 2026-08-27 — Demonstration computed on the one per-item outcome release found by the bounded search: union, all-miss, and residual coverage for the five specialized supervisors in BELLS 2025's released 170-prompt subset, registered as claim MC-002 (file bound by commit and sha256, reproduction script in CI) and rendered on the disclosure page. Census counts are unchanged — MC-002 is this record's own computation, not something the examined artifact printed.
- 2026-08-27 — Fresh-context adversarial verification, run same day and before anything was published, independently reproduced N/M/K = 19/13/4 and every MC-002 count, then found: (1) the wainjectbench-2025 reason line falsely called its ensemble members "academic and open-source" — Ensemble-I includes GPT-4o-Prompt, a commercial closed API; the row is corrected, and claim MC-001's commercial-API clause is withdrawn rather than reinterpreted, exactly as its forbidden rescues require. (2) A generator bug was deleting the letter "n" from the rendered criteria and exclusion notes on the public page (a malformed regex character class) — fixed, with the second exclusion rule's YAML typing corrected and the verifier extended to type-check exclusion rules. (3) The frozen phrase "separately attributable" had not said whether a product name was required. Rather than calling a post-review clarification pre-frozen, the record now exposes both readings: the primary treatment counts consistently distinguished, anonymized systems as separately attributable; the named-products-only sensitivity excludes unit42 and mechanically yields N/M/K = 18/12/4. Census schema v2 carries that executed sensitivity and a frozen-wording lock; criteria v1 itself is unchanged. (4) The IBM row's same-items evidence is marked as paraphrase. The primary counts remain N=19, M=13, K=4.
- 2026-08-28 — M ontology correction, made before any merge to main. The single number M was doing work it had not earned: the frozen criteria define "comparable" as shared items plus a shared event definition only, but a reader meets that word expecting matched operating thresholds and full exposure. compute_counts now derives an M ladder mechanically from fields already recorded on every row — 13 shared basis, 12 with no stated threshold mismatch (ML6 states the mismatch), 0 documenting matched thresholds with full exposure — and the page renders all three rungs. The proposition template gained an inline gloss saying which reading M uses. No row's evidence or classification changed; this is a weakening of what the headline implies, not a rescue of it. The strongest rung being 0 is the honest headline result and is now printed as such.
- 2026-08-28 — Adversarial release audit found that the repair still used the word "comparable" in the primary proposition and described the criteria lock as though it proved pre-search timing. Both claims were too strong. The public proposition now names only the mechanical shared-item/common-event basis, calls the four a heterogeneous joint-evidence discovery count rather than one estimand, and says precisely what the repository-history lock proves. No row, classification, or count changed; this is a further narrowing of language before any merge or public post.
What a correction can change
A correction can change a row, a count, or the scope of a claim. It cannot be used to retroactively narrow a criterion merely to preserve a preferred result. The census's claim envelope and forbidden-rescue rules are public in MC-001.