Recomputing from the committed observation rows yields any quantity different from the expected block beyond 1e-12, or a quoted prediction in the bound run report is shown to have been edited after the outcome was visible, or E3B is represented anywhere in this repository as a re-run, correction, or replacement of E3.
Consequence REJECTThis is the condition and consequence recorded in the registry. The vocabulary this is published in defines what each consequence commits the author to.
Exactly the 2,400 rows committed at experiments/e3b/results/observations.jsonl, produced 2026-09-06 by the same two classifiers at thresholds frozen in e3b_config.json (sha256 252a5db9…) before any injection item was scored. One pool (Lakera gandalf_ignore_instructions at 04737b65, MIT), one operating point each, static full exposure. This is a new experiment with its own freeze, not a re-run of E3: E3's result stands whatever this shows, and both are registered. scripts/verify_e3.py recomputes every quantity below from the committed rows alone. The guard that caught 400 of 400 has two live explanations — generalisation and near-duplicate leakage from an undisclosed corpus — and nothing in this run separates them.
Repairs declared unavailable in advance; using one after a failure would breach the recorded commitment.
- do not re-threshold, re-draw, or re-pool after the outcome and report the new numbers as this run
- do not present E3B as a repair of E3, or E3 as superseded by it
- do not restate either failed prediction, narrow it, or drop it from the record
- do not treat the zero-width interval as a measured absence of dependence
- a zero-width identified set means the marginals already fixed the joint miss; it is not a measurement that the two guards fail independently
- G1 catching 400 of 400 is not evidence that it is a good guardrail, and contamination is unverified rather than excluded
- G2 missing 159 of 400 is not evidence that it is a bad one
- 2,400 further rows prove the instrument runs; on this pool the question was degenerate, so they measure almost nothing
- the sharpening these two pilots suggest — that marginal-only reporting is uninformative in the middle of the marginal range and fully informative at its extremes — is a hypothesis and an engineering design gate, not a result of either pilot; neither run tested it
- no vendor, product, deployed stack, or population is described