Try it — three experiments, each with its command, expected result, and falsifiera result, not a compliment
DON'T TRUST THE GRAPHIC. REPRODUCE IT.
Do published per-guard scores determine what the stack misses together? Below: a 60-second proof, a 3-minute reconstruction from a public file, and a 15-minute audit you can run on an evaluation you know. Each says what it needs, what it should print, what would prove it wrong, and what it does not claim — before you run anything.
Thirty seconds, sound off
Same Scores, Different Worlds. Two guards that each miss 10 of 100 items can jointly miss anywhere from 0 to 10 of the 100 items; the two scores never move while the joint count runs the whole interval, and multiplying them selects one world.
Every number in the film is read from the claim registry; the film's manifest and render receipt are in films/same-scores-different-worlds/. Don't trust the animation. Run it: experiment A is the same construction as a script.
PROVED — Fréchet–Hoeffding bounds; both endpoint worlds are re-verified from a clean clone of the bound kernel in CI (CC-001, CC-004).
Falsifier
A joint distribution with marginals 0.10 and 0.10 whose both-miss probability lies outside [0.00, 0.10], or a witness the script prints that fails a marginal check, or the bound kernel returning other endpoints.
Non-claim
Nothing here measures a real guardrail; the population is a construction and no point inside the interval is selected. Independence is what one assumption chooses, not what the scores imply.
Film
Same Scores, Different Worlds — Two guards that each miss 10 of 100 items can jointly miss anywhere from 0 to 10 of the 100 items; the two scores never move while the joint count runs the whole interval, and multiplying them selects one world.
From a public per-item release, how many of 82 harmful prompts did all five supervisors miss — and what did the five published miss rates alone permit?
Needs
Python 3.10 or newer, `pip install -r requirements.txt` (PyYAML only), and network access to fetch one pinned CSV from GitHub. The file is hash-verified before a single count is taken.
Run
git clone https://github.com/Cubits11/cubits11.github.io.git && cd cubits11.github.io
python3 scripts/reanalyze_bells_subset.py
Variant
python3 scripts/identification.py --bells # the identification envelope, offline
Expected
final line: MC-002 reproduced: the missing column, computed from the bound public release, matches the registered claim.
OBSERVED counts on one released file (hash-verified, pinned commit); PROVED identified set from the marginals. Counting arithmetic, not a population estimate.
Falsifier
Recomputing from the pinned, hash-verified file yields any per-guard count, union, or all-miss different from the registered expected block; or the file at the pinned ref no longer matches its recorded sha256.
Non-claim
One author-selected 82-prompt stratum at unstated default configurations: not a vendor ranking, not a population estimate, not an observed deployed stack, and nothing about the adversarial prompts that have no per-item release.
Film
Thirteen Worlds, One File — The five published miss rates of the BELLS supervisors pin the all-miss count only to {0 … 12} of 82 — thirteen worlds; a same-denominator union fixes it at 9; the independence plug-in (3.49%, equivalent to 2.87 prompts in expectation) is a model's point, not an observed integer count.
Does an evaluation you know preserve what its compared systems miss together, or only what each misses alone?
Needs
Python 3.10 or newer, this repository, and one public evaluation of two or more guardrails that you can read. The script asks the census's frozen questions and classifies from your answers; it never reads the source for you.
Run
git clone https://github.com/Cubits11/cubits11.github.io.git && cd cubits11.github.io
python3 scripts/try_audit.py
Variant
python3 scripts/try_audit.py --answers fixtures/try/audit-example.json # a worked example, no prompts
Expected
final line: AUDIT TRY-C example-eval-2026 PRESENT printed_full_stack rung=threshold_not_contradicted (worked example, not a census row)
PROCEDURE — the same frozen criteria (v1) and predicates the census applies; your answers are OBSERVED by you, the classification is DERIVED from them, and a candidate row is a proposal until reviewed.
Falsifier
An evaluation that meets criteria v1, was public on or before 2026-08-27, and is absent from the census rejects the current count; an examined row whose source contradicts its recorded fields is corrected in the revision history.
Non-claim
A classification is a reporting fact about the artifact, never a performance finding about any system; a single audit by one person is a first-pass reading, exactly as the census's own rows are.
Film
The Missing Column, Counted — Of 20 public guardrail evaluations meeting frozen criteria, 14 report no joint-evidence artifact and 0 document matched operating thresholds with full exposure; the count is a fact about reporting inside a bounded search, corrected once in public and declared with its sensitivity.
QUALIFIED OUTCOMES — work done by someone who is not the author · bound to distribution/outcomes.yaml
independent reproductions
0
source corrections
0
paired outcome releases
0
upstream prs
0
human cold runs
0
Diagnostics, kept apart and never counted as outcomes: technical interactions 3; blinded comprehension trials 0. Zero is the recorded value. Be the first independent rerun → experiment A.
Report what you found
Prefilled forms, no vocabulary required. A different result is the most useful thing you can send; it is placed beside the claim it disagrees with, dated, and credited if you want it to be.
This is an open research problem with a correction path, not a request for help. The census examined 20 public evaluations: 5 preserve a joint-evidence artifact, 14 report none, and 0 document matched operating thresholds with full exposure. Each of the following would update the programme and be recorded with the prominence of the claim it changes:
an evaluation that meets the frozen criteria, was public on or before 2026-08-27, and is missing — the count changes (it has changed once already);
a row whose source contradicts its recorded fields — the row is corrected, dated, in the public file;
joint evidence for any examined evaluation — union, all-miss, leave-one-out, or aligned per-item outcomes — the row moves to PRESENT;
a narrower identified region for a registered bound, or a joint law that escapes one — the claim is narrowed or rejected under its registered falsifier;
a supposed missing column that is already present — the strongest kind of correction.