Pranav Bhave
Claim MC-004 — 11 of 21 in the registry Supported within scope

MC-004Warrant reviewed 2026-08-31 · expires 2026-10-30

Multimodal Safeguard Bench, recomputed from per-item verdicts

On the per-item guard-verdict files publicly released with Multimodal Safeguard Bench's full_run (pinned commit below; six verdict files and the release's two printed-metric files hash-verified), the three guards — Llama Guard 4, Llama Guard 3 Vision, ShieldGemma 2 — have harness-normalized native `unsafe` labels whose static Boolean OR is 1 on 192 of 200 harmful-labelled text items (8 all-zero rows, 4.0%) and all 200 harmful-labelled image items (0 all-zero rows), while the same OR is 1 on 32 of 250 benign-labelled text items (12.8%) and every one of 250 benign-labelled image items. Llama Guard 3 Vision alone has a 1-bit on all 250 released benign image items. The complete 2^3 native-label pattern table, leave-one-out bit ORs, and unique bit contributions for each released stratum are bound below. The release prints per-guard metrics and two two-guard compositions; it prints no three-guard bit OR, no all-zero-row count, no leave-one-out table, and no pattern decomposition — and every recomputed quantity that overlaps what it does print is asserted equal to the printed value.

01Falsifier — what changes this claim

Recomputing from the bound, hash-verified files yields any count different from the expected block, or any pinned file no longer matches its recorded sha256, or any recomputed quantity that overlaps the release's printed metrics disagrees with the printed value, or the committed blocked columns are shown not to be the named guards' verdicts.

Consequence REJECT

This is the condition and consequence recorded in the registry. The vocabulary this is published in defines what each consequence commits the author to.

02Scope

Exactly the six committed verdict files of the full_run directory at the bound commit: 400 harmful and 500 benign items, each stratum half text and half image, one Boolean `blocked` adapter bit per item per guard. Let L be the release's harmful/benign label, A_s each guard's native `unsafe` predicate, B_s the harness mapping of that predicate to `blocked`, and O = OR_s B_s. The bound arithmetic reports B and O conditional on L. Because the harness uses `blocked` items to suppress target generation, O is also a valid counterfactual harness-block decision for these released rows under a fixed block-on-any rule. It does not establish a source-defined translation from every A_s to one shared catch event E: Llama Guard 3 Vision is a multimodal prompt/response classifier, while ShieldGemma 2 is an image-only three-policy classifier, and the pipeline merely maps each native `unsafe` label to `blocked`. Counting arithmetic only, at the guards' released native operating rules (the Llama guards emit labels autoregressively; ShieldGemma 2 blocks above a 0.5 policy-violation probability; the source documents no matched operating-point calibration, so M-ladder threshold comparability is unchanged). ShieldGemma 2's text-item bits are deterministic passes as committed and documented upstream, which is why its text column is zero — an explicit released outcome, not missing data. Strata are computed exactly as released and never folded; file totals are sums of the bound strata and carry no additional content. Computed by scripts/reanalyze_msbench.py from scripts/mjgd_reference.py; the expected block below is what CI re-asserts against the downloaded, hash-verified files.

03Forbidden rescues

Repairs declared unavailable in advance; using one after a failure would breach the recorded commitment.

04Non-claims — what this does not license
05Binding and freshness
Binding
PatrickKollman/Multimodal-Safeguard-Bench/tree/fb6f32e6b50b6faad833b815ebcd80afd2068bff/results/full_run @ fb6f32e6
Six guard-verdict JSONL files are the data inputs; metrics.json and ensemble.json are additionally pinned because the reproduction script asserts agreement with every overlapping printed value. All eight sha256 hashes are recorded in scripts/reanalyze_msbench.py and verified before a single count is taken.
Reviewed
2026-08-31 · window 60 days
Expires
2026-10-30 — after this date the recorded review is overdue; this does not make the claim false
Triggers
  • executable fires when a bound verdict file on the default branch changes (6 bound sources)
  • executable fires when the release's printed per-guard metrics change
  • executable fires when the release's printed composition metrics change
  • executable fires when the guards' native adapter semantics change
  • executable fires when the harness label-to-block action changes
  • executable fires when the stated route or operating-point semantics change
  • manual the authors publish a corrected or extended full_run, print their own three-guard joint statistics, document a matched operating-point calibration, or publish a source-defined translation from all native predicates to one shared catch event
Dimensions
visibilityPublicprovenanceMachine-generated, owner-executedsupport roleExecuted outputmaturityExperimental