The argument in text
For marginals (0.10, 0.10) the both-fail probability is identified only to [0.0, 0.10] and independence names 0.01 (CC-001, CC-004). As of 2026-08-27, of 20 examined evaluations 14 establish a shared basis, 0 document matched thresholds with full exposure, and 5 provide a joint-evidence artifact, 2 of them via a per-item release (MC-001). On the BELLS 2025 released subset the five marginals identify the all-miss count only to {0 … 12} of 82; the released rows give 9; the independence plug-in expects 2.87 (MC-002, MC-003).
A montage of four bound facts with no new arithmetic: the construction and interval are the Same Scores master's (illustrative population, CC-001 input, CC-004 witnesses); the census counts are verify_census.compute_counts on census.yaml at the freeze, one primary reviewer, bounded search; the BELLS numbers are counting arithmetic on one author-selected file at native, unstated operating points. Nothing on screen is a measured deployed stack.
What this does not establish
- does not certify any stacked system as safe and shows no measured deployed stack
- does not estimate or recover the true dependence between any two guards; the amber points are what one assumption would select
- does not claim any evaluated guardrail performs poorly; an unfilled column is a reporting fact
- does not treat the BELLS 9/82 as a rate or a population estimate; it is counting arithmetic on one author-selected released file
- the census is a bounded single-reviewer search, not proof of universal absence
- no new mathematics is claimed; the bounds are Fréchet's and the lower endpoint is Bonferroni's
- not a claim that guardrails in general fail together — this stratum cannot separate shared blind spots from prompt-difficulty heterogeneity
- not a vendor evaluation; the released verdicts are at unstated default configurations, and one supervisor fires exactly once in the 170 released rows
- the identified interval is what the marginals leave open, not a prediction about any deployed stack
- leave-one-out unions identify only exclusive full-stack coverage; they do not identify pairwise or higher-order overlap, Shapley values, or causal contribution
- says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release
- the subset is author-selected with an unstated rule — nothing here estimates any system's true rate, and no confidence interval is offered because the sampled population is undefined
- not a ranking or endorsement of any vendor; the released verdicts are at unstated default configurations
- says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release
- the release-recomputed-to-plug-in ratio describes this subset's arithmetic, not a general law of guardrail dependence
- not an observed deployed five-guard stack; interpreting this static OR aggregation as a stack requires the separate full-exposure, parallel, fixed-operating-point assumptions
- does not claim any evaluated guardrail stack performs poorly; an unfilled column is a reporting fact, not a performance finding
- does not claim the unmeasured joint statistics would reveal dependence; measuring instead of assuming is the point
- does not audit the quality of any per-system evaluation beyond the fields each row records
- does not claim exhaustive coverage beyond the corrected, examined artifact set or coverage after the stated date
- does not certify any stacked system as safe
- does not recover or estimate the unknown dependence
- does not select a point inside the returned interval
What would require a correction
A clean-clone execution of the bound kernel on marginals (0.10, 0.10) returns AND bounds other than [0.0, 0.10] (CC-001/CC-004, REJECT); scripts/verify_census.py --counts prints counts other than 20 examined, 14 with a shared basis, 0 with matched thresholds and full exposure, 5 with a joint-evidence artifact, and 2 with per-item releases (MC-001, REJECT); scripts/reanalyze_bells_subset.py recomputes an all-miss other than 9/82 or scripts/identification.py a set other than {0 … 12} (MC-002/MC-003, REJECT); or any rendered frame shows a readout other than 10/100 for either guard while the BOTH counter moves — in each case bind_facts.py --check fails before this film can be re-rendered.
Read the supporting record: CC-001 · CC-004 · MC-001 · MC-002 · MC-003.
Sources and render record
Film manifest · Render receipt · First site deployment record
Current render: 2026-09-15. The receipt records the rendered inputs and output hash. It checks provenance, not scientific validity.