Pranav Bhave

Film · 34 seconds · silent

Leave One Out

Marginal catch counts plus the union only bound what each guard alone catches; the released leave-one-out unions identify it — 18, 3, 0, 0, 0 — so three of five supervisors add no exclusive coverage on this stratum.

Read the argumentDownload MP4First published here 2026-09-02

The argument in text

On the 82 prompts labelled harmful in the released BELLS subset, the OR of released verdicts flags 73; removing NeMo Guardrails leaves 55 flagged (exclusive coverage 18), removing Lakera Guard leaves 70 (3), and removing LangKit, Prompt Guard, or LLM Guard leaves 73 (0). From catch counts and the union alone those exclusive counts are only bounded: NeMo 0–21, Lakera 0–3, LangKit 0–3, Prompt Guard 0–3, LLM Guard 0.

Exactly the released file at the bound commit; counting arithmetic at the vendors' released binary verdicts. The lamp counts, union, and leave-one-out unions are registered values re-asserted against the hash-verified file in CI (MC-002); the bounds are MC-003's. Which windows each lamp lights is a construction satisfying every registered count and all five leave-one-out unions; the released rows are never redistributed, and pairwise overlaps are neither registered nor claimed.

What this does not establish

  • not a vendor ranking, a causal attribution, or evidence about any product outside this stratum
  • the released verdicts are at unstated default configurations; one supervisor fires exactly once in the released rows
  • leave-one-out unions identify exclusive full-stack coverage only; not pairwise or higher-order overlap, Shapley values, or ensemble structure
  • window positions are a construction; the released rows are not redistributed and the true overlaps beyond the registered aggregates are not shown
  • says nothing about adversarial prompts, which dominate the evaluation and have no per-item release
  • the subset is author-selected with an unstated rule — nothing here estimates any system's true rate, and no confidence interval is offered because the sampled population is undefined
  • not a ranking or endorsement of any vendor; the released verdicts are at unstated default configurations
  • says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release
  • the release-recomputed-to-plug-in ratio describes this subset's arithmetic, not a general law of guardrail dependence
  • not an observed deployed five-guard stack; interpreting this static OR aggregation as a stack requires the separate full-exposure, parallel, fixed-operating-point assumptions
  • no new mathematics is claimed; the bounds are Fréchet's and the lower endpoint is Bonferroni's
  • not a claim that guardrails in general fail together — this stratum cannot separate shared blind spots from prompt-difficulty heterogeneity
  • not a vendor evaluation; the released verdicts are at unstated default configurations, and one supervisor fires exactly once in the 170 released rows
  • the identified interval is what the marginals leave open, not a prediction about any deployed stack
  • leave-one-out unions identify only exclusive full-stack coverage; they do not identify pairwise or higher-order overlap, Shapley values, or causal contribution
  • says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release

What would require a correction

Recomputing from the bound, hash-verified file yields any per-guard catch, union, or leave-one-out union different from the expected block (MC-002, REJECT); or an exact finite catch-set arrangement lies outside the stated exclusive-coverage bounds (MC-003, REJECT); or the constructed wall fails any registered count — bind_facts.py asserts every constraint before the film can read it.

Read the supporting record: MC-002 · MC-003.

Sources and render record

Film manifest · Render receipt · First site deployment record

Current render: 2026-09-15. The receipt records the rendered inputs and output hash. It checks provenance, not scientific validity.