Pranav Bhave
Claim E3B-001 — 19 of 21 in the registry Supported within scope

E3B-001Warrant reviewed 2026-09-08 · expires 2027-01-06

The same classifiers on the attack family they were built for

E3B put the same two classifiers on the attack family they were built for — 400 real prompt injections — and produced a further 2,400 committed observation rows. Its primary prediction FAILED and so did the prediction the redraw existed to test. One guard missed nothing (0.0000, catching 400 of 400) and the other missed 0.3975, so the Fréchet interval the marginals allow is [0.0000, 0.0000] — zero points wide — the observed joint miss is exactly 0.0000, and the bootstrap CI is the degenerate [0, 0]. E3B was pre-registered to produce an interval wider than 10 percentage points; it produced one narrower than E3's. The prediction that the observed joint miss would lie inside the interval HELD, trivially.

01Falsifier — what changes this claim

Recomputing from the committed observation rows yields any quantity different from the expected block beyond 1e-12, or a quoted prediction in the bound run report is shown to have been edited after the outcome was visible, or E3B is represented anywhere in this repository as a re-run, correction, or replacement of E3.

Consequence REJECT

This is the condition and consequence recorded in the registry. The vocabulary this is published in defines what each consequence commits the author to.

02Scope

Exactly the 2,400 rows committed at experiments/e3b/results/observations.jsonl, produced 2026-09-06 by the same two classifiers at thresholds frozen in e3b_config.json (sha256 252a5db9…) before any injection item was scored. One pool (Lakera gandalf_ignore_instructions at 04737b65, MIT), one operating point each, static full exposure. This is a new experiment with its own freeze, not a re-run of E3: E3's result stands whatever this shows, and both are registered. scripts/verify_e3.py recomputes every quantity below from the committed rows alone. The guard that caught 400 of 400 has two live explanations — generalisation and near-duplicate leakage from an undisclosed corpus — and nothing in this run separates them.

03Forbidden rescues

Repairs declared unavailable in advance; using one after a failure would breach the recorded commitment.

04Non-claims — what this does not license
05Binding and freshness
Binding
observations.jsonl
The support is the measurement itself: 2,400 further per-item, per-guard rows. RESULT.md is the interpretation, hash-pinned by a trigger below. scripts/verify_e3.py re-derives every registered number from the rows alone on every push.
Reviewed
2026-09-08 · window 120 days
Expires
2027-01-06 — after this date the recorded review is overdue; this does not make the claim false
Triggers
  • executable fires when a committed observation row changes without a registry re-review
  • executable fires when the recorded run result changes without a registry re-review
  • executable fires when the run report changes, including any edit to a quoted prediction
  • manual the contamination status of either guard's training corpus becomes checkable
Dimensions
visibilityPublicprovenanceMachine-generated, owner-executedsupport roleExecuted outputmaturityExperimental