Pranav Bhave

Film · 36 seconds · silent

Thirteen Worlds, One File

The five published miss rates of the BELLS supervisors pin the all-miss count only to {0 … 12} of 82 — thirteen worlds; a same-denominator union fixes it at 9; the independence plug-in (3.49%, equivalent to 2.87 prompts in expectation) is a model's point, not an observed integer count.

Read the argumentDownload MP4First published here 2026-09-02

The argument in text

On the 82 prompts labelled harmful in the released BELLS subset, the five per-guard miss counts (12, 30, 57, 77, 82) identify the static all-miss count only up to {0/82 … 12/82}; the released OR-union flags 73, so the all-miss count is 9 (11.0%); the independence plug-in is 3.49%, or 2.87 prompts, and the recomputed rate is 3.14× that plug-in on this file.

Exactly the released file at the bound commit, counting arithmetic at the vendors' released binary verdicts. The bead COUNTS per rod, the union, and the all-miss count are registered values re-asserted against the hash-verified file in CI (MC-002). The bead POSITIONS are a construction: every arrangement shown satisfies the five counts, and the final one also satisfies the union; the released rows are never redistributed. No population estimate, no vendor ranking, no claim about adversarial prompts.

What this does not establish

  • not an observed deployed five-guard stack; the OR of released columns is a static aggregation
  • not a population estimate and not a confidence interval; the subset was author-selected under an unstated rule
  • not a ranking or endorsement of any vendor; verdicts are at unstated default configurations
  • the bead positions are not the released rows; pairwise overlaps are not identified by the aggregates shown and are not claimed
  • the 3.14× ratio describes this subset's arithmetic, not a law of guardrail dependence
  • no new mathematics is claimed; the bounds are Fréchet's and the lower endpoint is Bonferroni's
  • not a claim that guardrails in general fail together — this stratum cannot separate shared blind spots from prompt-difficulty heterogeneity
  • not a vendor evaluation; the released verdicts are at unstated default configurations, and one supervisor fires exactly once in the 170 released rows
  • the identified interval is what the marginals leave open, not a prediction about any deployed stack
  • leave-one-out unions identify only exclusive full-stack coverage; they do not identify pairwise or higher-order overlap, Shapley values, or causal contribution
  • says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release
  • the subset is author-selected with an unstated rule — nothing here estimates any system's true rate, and no confidence interval is offered because the sampled population is undefined
  • not a ranking or endorsement of any vendor; the released verdicts are at unstated default configurations
  • says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release
  • the release-recomputed-to-plug-in ratio describes this subset's arithmetic, not a general law of guardrail dependence
  • not an observed deployed five-guard stack; interpreting this static OR aggregation as a stack requires the separate full-exposure, parallel, fixed-operating-point assumptions

What would require a correction

Recomputing from the bound, hash-verified file yields any per-guard catch count, union, or all-miss different from the expected block (MC-002, REJECT); or a joint law with the five stated miss rates attains an all-miss rate outside [0, 12/82] (MC-003, REJECT); or the film's sliding rod ever produces an all-miss count outside {0 … 12} — checkable by re-rendering and reading the counter at every frame.

Read the supporting record: MC-002 · MC-003.

Sources and render record

Film manifest · Render receipt · First site deployment record

Current render: 2026-09-15. The receipt records the rendered inputs and output hash. It checks provenance, not scientific validity.