The argument in text
On the 82 prompts labelled harmful in the released BELLS subset, the five per-guard miss counts (12, 30, 57, 77, 82) identify the static all-miss count only up to {0/82 … 12/82}; the released OR-union flags 73, so the all-miss count is 9 (11.0%); the independence plug-in is 3.49%, or 2.87 prompts, and the recomputed rate is 3.14× that plug-in on this file.
Exactly the released file at the bound commit, counting arithmetic at the vendors' released binary verdicts. The bead COUNTS per rod, the union, and the all-miss count are registered values re-asserted against the hash-verified file in CI (MC-002). The bead POSITIONS are a construction: every arrangement shown satisfies the five counts, and the final one also satisfies the union; the released rows are never redistributed. No population estimate, no vendor ranking, no claim about adversarial prompts.
What this does not establish
- not an observed deployed five-guard stack; the OR of released columns is a static aggregation
- not a population estimate and not a confidence interval; the subset was author-selected under an unstated rule
- not a ranking or endorsement of any vendor; verdicts are at unstated default configurations
- the bead positions are not the released rows; pairwise overlaps are not identified by the aggregates shown and are not claimed
- the 3.14× ratio describes this subset's arithmetic, not a law of guardrail dependence
- no new mathematics is claimed; the bounds are Fréchet's and the lower endpoint is Bonferroni's
- not a claim that guardrails in general fail together — this stratum cannot separate shared blind spots from prompt-difficulty heterogeneity
- not a vendor evaluation; the released verdicts are at unstated default configurations, and one supervisor fires exactly once in the 170 released rows
- the identified interval is what the marginals leave open, not a prediction about any deployed stack
- leave-one-out unions identify only exclusive full-stack coverage; they do not identify pairwise or higher-order overlap, Shapley values, or causal contribution
- says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release
- the subset is author-selected with an unstated rule — nothing here estimates any system's true rate, and no confidence interval is offered because the sampled population is undefined
- not a ranking or endorsement of any vendor; the released verdicts are at unstated default configurations
- says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release
- the release-recomputed-to-plug-in ratio describes this subset's arithmetic, not a general law of guardrail dependence
- not an observed deployed five-guard stack; interpreting this static OR aggregation as a stack requires the separate full-exposure, parallel, fixed-operating-point assumptions
What would require a correction
Recomputing from the bound, hash-verified file yields any per-guard catch count, union, or all-miss different from the expected block (MC-002, REJECT); or a joint law with the five stated miss rates attains an all-miss rate outside [0, 12/82] (MC-003, REJECT); or the film's sliding rod ever produces an all-miss count outside {0 … 12} — checkable by re-rendering and reading the counter at every frame.
Sources and render record
Film manifest · Render receipt · First site deployment record
Current render: 2026-09-15. The receipt records the rendered inputs and output hash. It checks provenance, not scientific validity.