A qualifying public evaluation meeting criteria v1 and published on or before 2026-08-27 is shown to be absent from this corrected census, or an examined row is shown to misreport its source such that recomputed N, M, or K differ from the stated 20/14/5.
Consequence REJECTThis is the condition and consequence recorded in the registry. The vocabulary this is published in defines what each consequence commits the author to.
A census claim about public reporting, bounded by the documented search protocol documented in census.yaml. Its literal inclusion wording is locked in repository history before row classification; that is a reproducibility lock, not an independent preregistration. It covers the artifacts that search found, not everything in existence; it says nothing about how well any evaluated system or stack performs, and nothing about what the unmeasured joint statistics would show. N, M, and K are recomputed mechanically from census.yaml by scripts/verify_census.py; this envelope's expected block is cross-checked against that computation in CI. Since census schema v3 the joint-evidence mode counts in this note are derived the same way, from a declared joint_scope_additional field, rather than counted by hand here; scripts/verify_facts.py binds every current census numeral on every public page, and on the deployed page after release, to the quantity it asserts. M is a ladder, not a comparability verdict: 14 document shared items and a common event definition, 12 have no stated threshold mismatch, and 0 document matched operating thresholds together with full exposure. Nothing here asserts that the 14 are interchangeable at a common operating point; the opposite is the finding. The 5 is an inclusive discovery count of noninterchangeable artifacts, not an all-miss rate or a deployment conclusion. Four artifacts print at least one composition result and two release aligned per-item outcomes from which joint statistics are directly computable; one artifact does both, so those descriptions intentionally overlap. The frozen phrase "separately attributable" did not specify whether a product or model name must be printed: the primary count treats systems a source distinguishes and reports consistently as separately attributable even when anonymized; the declared named-products-only sensitivity excludes Unit42 and mechanically yields N/M/K = 19/13/5. A drafted further clause — that no printed joint statistic among the then-four covers a commercial guardrail API — was withdrawn before publication when same-day adversarial review produced a live counterexample under one defensible reading (WAInjectBench's printed ensemble includes GPT-4o prompted as a detector); this claim's forbidden rescues bar narrowing "commercial" after the fact, so the clause was dropped rather than reinterpreted, and the event is recorded in the census revision history. On 2026-08-30, a post-release source audit found the distinct Multimodal Safeguard Bench repository, whose currently public Git history carries a pre-cutoff commit timestamp, and which the documented search had missed. That met this claim's prior REJECT falsifier: the 19/13/4 envelope is rejected and retained in census revision history. This 20/14/5 envelope is its corrected, superseding proposition, not a reinterpretation of the former count.
Repairs declared unavailable in advance; using one after a failure would breach the recorded commitment.
- do not reinterpret "joint statistic" after a counterexample appears in order to keep the stated counts
- do not reclassify an examined row without recording the change and its reason in the census correction history
- do not treat the bounded-search disclaimer as license to ignore a demonstrated miss — a miss rejects the current counts and the corrected census must state new ones
- do not cite readership, links, or reuse of the census as evidence of its accuracy
- does not claim any evaluated guardrail stack performs poorly; an unfilled column is a reporting fact, not a performance finding
- does not claim the unmeasured joint statistics would reveal dependence; measuring instead of assuming is the point
- does not audit the quality of any per-system evaluation beyond the fields each row records
- does not claim exhaustive coverage beyond the corrected, examined artifact set or coverage after the stated date