Pranav Bhave
Claim MC-001 — 8 of 21 in the registry Supported within scope

MC-001Warrant reviewed 2026-09-14 · expires 2026-11-13

Of 20 guardrail evaluations, 5 with joint-evidence artifacts

As of 2026-08-27, among 20 public guardrail evaluations meeting the Missing Column Census's frozen inclusion criteria (v1), 14 establish a shared item set and a common event definition; 5 provide one of the census's declared joint-evidence artifacts: a printed composition result or per-item outcomes that make one directly computable.

01Falsifier — what changes this claim

A qualifying public evaluation meeting criteria v1 and published on or before 2026-08-27 is shown to be absent from this corrected census, or an examined row is shown to misreport its source such that recomputed N, M, or K differ from the stated 20/14/5.

Consequence REJECT

This is the condition and consequence recorded in the registry. The vocabulary this is published in defines what each consequence commits the author to.

02Scope

A census claim about public reporting, bounded by the documented search protocol documented in census.yaml. Its literal inclusion wording is locked in repository history before row classification; that is a reproducibility lock, not an independent preregistration. It covers the artifacts that search found, not everything in existence; it says nothing about how well any evaluated system or stack performs, and nothing about what the unmeasured joint statistics would show. N, M, and K are recomputed mechanically from census.yaml by scripts/verify_census.py; this envelope's expected block is cross-checked against that computation in CI. Since census schema v3 the joint-evidence mode counts in this note are derived the same way, from a declared joint_scope_additional field, rather than counted by hand here; scripts/verify_facts.py binds every current census numeral on every public page, and on the deployed page after release, to the quantity it asserts. M is a ladder, not a comparability verdict: 14 document shared items and a common event definition, 12 have no stated threshold mismatch, and 0 document matched operating thresholds together with full exposure. Nothing here asserts that the 14 are interchangeable at a common operating point; the opposite is the finding. The 5 is an inclusive discovery count of noninterchangeable artifacts, not an all-miss rate or a deployment conclusion. Four artifacts print at least one composition result and two release aligned per-item outcomes from which joint statistics are directly computable; one artifact does both, so those descriptions intentionally overlap. The frozen phrase "separately attributable" did not specify whether a product or model name must be printed: the primary count treats systems a source distinguishes and reports consistently as separately attributable even when anonymized; the declared named-products-only sensitivity excludes Unit42 and mechanically yields N/M/K = 19/13/5. A drafted further clause — that no printed joint statistic among the then-four covers a commercial guardrail API — was withdrawn before publication when same-day adversarial review produced a live counterexample under one defensible reading (WAInjectBench's printed ensemble includes GPT-4o prompted as a detector); this claim's forbidden rescues bar narrowing "commercial" after the fact, so the clause was dropped rather than reinterpreted, and the event is recorded in the census revision history. On 2026-08-30, a post-release source audit found the distinct Multimodal Safeguard Bench repository, whose currently public Git history carries a pre-cutoff commit timestamp, and which the documented search had missed. That met this claim's prior REJECT falsifier: the 19/13/4 envelope is rejected and retained in census revision history. This 20/14/5 envelope is its corrected, superseding proposition, not a reinterpretation of the former count.

03Forbidden rescues

Repairs declared unavailable in advance; using one after a failure would breach the recorded commitment.

04Non-claims — what this does not license
05Binding and freshness
Binding
census.yaml
Mutable link by design; the local-content trigger below detects edits, and every row in the file binds to its own primary source with passages, dates, and a correction history.
Reviewed
2026-09-14 · window 60 days
Expires
2026-11-13 — after this date the recorded review is overdue; this does not make the claim false
Triggers
  • executable fires when the census changes without a claim re-review; re-reviewed 2026-09-14 (America/New_York), confirming the owner report and independently comparing the bytes: the only change from the previously pinned bytes is one appended AgentDrift / DriftNet entry under unexamined_candidates. All classified content is identical; verify_census.py confirms N/M/K 20/14/5 and M ladder 14/12/0. This hash review does not refresh any row-level source check. re-reviewed 2026-09-01 after declaring the pre-freeze-visibility-evidenced sensitivity: the row that moved 19/13/4 to 20/14/5 is in the 20 on GitHub-recorded dates that do not record when its repository became publicly visible; the Wayback Machine's CDX index held no capture of it at any date when checked on 2026-09-01, and no other archive was queried. The sensitivity's 19/13/4 alternative envelope is now asserted by verify_census.py rather than living in revision prose. No row, classification, evidence field, or primary count changed; N/M/K remain 20/14/5 and the M ladder remains 14/12/0.
  • manual a new qualifying evaluation is published or reported, an author corrects a row, or a benchmark publishes a previously absent joint statistic
Dimensions
visibilityPublicprovenanceOwner-authored, publicsupport roleSite documentmaturityExperimental