A joint law with the stated marginals achieves an all-miss rate outside the interval; an exact finite catch-set arrangement lies outside the stated finite grid or exclusive-coverage bounds; or recomputing from the bound BELLS file yields any count different from the expected block.
Consequence REJECTThis is the condition and consequence recorded in the registry. The vocabulary this is published in defines what each consequence commits the author to.
The mathematics is classical — Bonferroni for the lower endpoint, monotonicity for the upper, Fréchet for sharpness — and nothing here claims to have derived it. What is contributed is the application: pricing, in probability units, what marginal-only guardrail reporting leaves undetermined, and identifying leave-one-out unions as a compact privacy-preserving disclosure of exclusive full-stack coverage. They do not identify pairwise or higher-order overlap, Shapley values, or causal structure. Every statement is conditional on block-on-any composition, per-item decisions that are functions of the item alone, a fixed population, full exposure, and a fixed operating point; it says nothing about sequential routes where downstream measurement is censored by upstream blocks, agentic systems where an intervention changes the trajectory, or any vendor's product. The empirical half is one stratum of one author-selected subset: 82 harmful prompts, with 50 benign and 38 borderline held as separate strata and never folded in. It is a BELLS-specific reproduction, not an assertion that BELLS is the census's only per-item outcome release. scripts/reanalyze_bells_subset.py re-computes the BELLS counts from the hash-verified upstream file in CI; scripts/identification.py checks the stated identification transformations against that registered envelope.
Repairs declared unavailable in advance; using one after a failure would breach the recorded commitment.
- do not change the composition rule after the fact to preserve the direction of the result
- do not restate "identified only up to" as "estimated to be" — a bound is not an estimate
- do not quote where a release-recomputed value sits inside the interval as a score or a percentage of a gap closed; the endpoints come from adversarial couplings with no detector-behavioural content
- do not cite the zero-exclusive-coverage results as a vendor ranking, a causal attribution, or evidence about any product outside this stratum
- do not fold the borderline stratum into either denominator to change any figure
- no new mathematics is claimed; the bounds are Fréchet's and the lower endpoint is Bonferroni's
- not a claim that guardrails in general fail together — this stratum cannot separate shared blind spots from prompt-difficulty heterogeneity
- not a vendor evaluation; the released verdicts are at unstated default configurations, and one supervisor fires exactly once in the 170 released rows
- the identified interval is what the marginals leave open, not a prediction about any deployed stack
- leave-one-out unions identify only exclusive full-stack coverage; they do not identify pairwise or higher-order overlap, Shapley values, or causal contribution
- says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release