The argument in text
As of 2026-08-27, among 20 examined public evaluations, 14 establish a shared item set and common event definition, 12 state no threshold mismatch, 0 document matched thresholds with full exposure; 5 provide a declared joint-evidence artifact (4 print a composition result, 2 release computable per-item outcomes, 1 does both), 14 report none, 1 is not comparable. The 19/13/4 envelope was rejected on 2026-08-30 by the claim's own falsifier and retained in the revision history; a declared sensitivity (archive-evidenced pre-freeze visibility) computes 19/13/4.
A census of reporting bounded by its documented search protocol, classified by a single primary reviewer with no second review yet. It says nothing about how any evaluated system or stack performs. Every count on screen is verify_census.compute_counts on census.yaml; the per-row ladder flags use the identical predicates and are asserted to reproduce 14/12/0.
What this does not establish
- does not claim any evaluated guardrail or stack performs poorly; an unfilled column is a reporting fact, not a performance finding
- does not claim the unmeasured joint statistics would reveal dependence
- does not audit the quality of any per-system evaluation; the per-system dashes carry no scores
- does not claim exhaustive coverage beyond the examined set; the 25 unexamined candidates are unknown, not absent
- the sensitivity envelope is never a bare replacement for 20/14/5
What would require a correction
A qualifying public evaluation meeting criteria v1 and published on or before 2026-08-27 is shown absent from the census, or a row misreports its source such that recomputed N, M, or K differ from 20/14/5 (MC-001, REJECT); or scripts/verify_census.py --counts prints counts other than the ones on screen — in which case bind_facts.py --check fails before this film can be re-rendered.
Read the supporting record: MC-001.
Sources and render record
Film manifest · Render receipt · First site deployment record
Current render: 2026-09-15. The receipt records the rendered inputs and output hash. It checks provenance, not scientific validity.