Recomputing from the bound, hash-verified files yields any count different from the expected block, or any pinned file no longer matches its recorded sha256, or any recomputed quantity that overlaps the release's printed metrics disagrees with the printed value, or the committed blocked columns are shown not to be the named guards' verdicts.
Consequence REJECTThis is the condition and consequence recorded in the registry. The vocabulary this is published in defines what each consequence commits the author to.
Exactly the six committed verdict files of the full_run directory at the bound commit: 400 harmful and 500 benign items, each stratum half text and half image, one Boolean `blocked` adapter bit per item per guard. Let L be the release's harmful/benign label, A_s each guard's native `unsafe` predicate, B_s the harness mapping of that predicate to `blocked`, and O = OR_s B_s. The bound arithmetic reports B and O conditional on L. Because the harness uses `blocked` items to suppress target generation, O is also a valid counterfactual harness-block decision for these released rows under a fixed block-on-any rule. It does not establish a source-defined translation from every A_s to one shared catch event E: Llama Guard 3 Vision is a multimodal prompt/response classifier, while ShieldGemma 2 is an image-only three-policy classifier, and the pipeline merely maps each native `unsafe` label to `blocked`. Counting arithmetic only, at the guards' released native operating rules (the Llama guards emit labels autoregressively; ShieldGemma 2 blocks above a 0.5 policy-violation probability; the source documents no matched operating-point calibration, so M-ladder threshold comparability is unchanged). ShieldGemma 2's text-item bits are deterministic passes as committed and documented upstream, which is why its text column is zero — an explicit released outcome, not missing data. Strata are computed exactly as released and never folded; file totals are sums of the bound strata and carry no additional content. Computed by scripts/reanalyze_msbench.py from scripts/mjgd_reference.py; the expected block below is what CI re-asserts against the downloaded, hash-verified files.
Repairs declared unavailable in advance; using one after a failure would breach the recorded commitment.
- do not fold text and image strata — or harmful and benign files — into pooled denominators to move any figure; the strata are bound exactly as released
- do not substitute verdicts from the release's other run directories (adaptive_run, carrier sweeps, rendering probes) to preserve a number bound to full_run
- do not reinterpret ShieldGemma 2's deterministic text passes as missing data to shrink a denominator after seeing the results
- do not cite the benign-image 250/250 union without attributing it to Llama Guard 3 Vision's 250/250 column, and do not cite the harmful strata without the benign strata
- do not call an OR of the harness-normalized native labels a shared-event catch union, stack safety result, or three-independent- guard finding unless a source-defined event translation is added
- do not recast these file counts as population estimates, with or without an interval, if the primary counts are challenged
- not a shared-event catch statistic: the common `blocked` bit is a harness normalization of distinct native predicates, and no source-defined translation to a common event E has been identified
- not a population estimate — the items derive from HarmBench and XSTest under the release's own construction, and no interval is offered because no sampled population is defined
- not a ranking, endorsement, or indictment of any guard or vendor; the verdicts are at native, unmatched operating rules
- not an observed deployed three-guard stack; interpreting the OR aggregation as a stack requires the separate full-exposure, parallel, fixed-operating-point assumptions
- says nothing about the release's carrier-prompt, adversarial-UAP, or cross-VLM runs, which are separate artifacts with their own contracts
- the counts are about the committed verdict bytes; the release's own changelog documents that ShieldGemma 2's image scores are sensitive to the text-rendering stack, so nothing here predicts what any guard would do under a different rendering environment
- the zero three-guard image all-zero-bit count is a counting fact about these 200 released image items, not evidence of general image attack safety