Pranav Bhave
Claim E7B-001 — 21 of 21 in the registry REJECTED AS STATED

E7B-001REJECTED AS STATED 2026-09-10Warrant reviewed 2026-09-10 · expires 2027-01-08

Rejected: calibration uses the shared pool for every judge

REJECTED AS STATED on 2026-09-10. Calibration uses the all-nine shared pool instead of each judge’s own pool, changing one threshold from 0.95 to 1.00. Pair evaluation also uses the global intersection. The preregistration-conforming confirmation claim is rejected; the original HELD labels are historical outputs. Historical assertion retained for audit, not an accepted current finding: E7B is the first non-degenerate preregistered measurement this repository has produced, and all five of its predictions held on data unread when they were fixed. Against a preregistration frozen at f646e136 before any hold-out byte was retrieved, six of nine published guardrail-judge score sets reached a 5% benign false-flag budget on their shared 96-item AgentDojo pool (21 injection goals, 75 benign) and yielded 15 non-degenerate pairs. Median excess joint miss was +0.2018 and mean +0.1852 (P1, P2 HELD); the median pair's observed joint miss sat 0.900 of the way to the Fréchet upper bound its own marginals allow (P3 HELD); every observed joint lay inside its interval (P4 HELD); and every one of the 15 pairs had an observed joint miss above the independence product (P5 HELD). On the modal pair two judges each miss about half the injection goals, independence predicts a 0.249 both-miss rate, and the observed rate is 0.476 — exactly the upper bound, meaning one judge's misses are a subset of the other's. Seven of the fifteen pairs sit at that bound.

01Falsifier — what changes this claim

Recomputing from the committed rows yields any registered quantity different beyond 1e-9, or the preregistration's commit is shown not to precede the results commit, or any of the nine hold-out score files is shown to have been retrieved or inspected before f646e136, or the operating-point rule applied in experiments/e7b/run/measure.py is shown to differ from the one PREREG.md states, or an observed joint miss is shown to lie outside the Fréchet interval its own marginals allow.

Consequence REJECT

This is the condition and consequence recorded in the registry. The vocabulary this is published in defines what each consequence commits the author to.

02Scope

Historical run only. Formal disposition: experiments/e7b/CORRECTION-2026-09-10.md. Frozen inputs, estimator and outputs are unchanged. The expected block below records original run values and must not be used as current confirmatory evidence. Original scope, including interpretations now rejected with this claim: Exactly the 15 rows committed at experiments/e7b/results/observations.jsonl, computed 2026-09-10 from nine published score files in shawnray-research/certified-agent-guardrails at commit 79097583be7786976ea1b9ae79f3ff900d9e66b7 (MIT). Labels are the 35 strings of _inj_goals.json in that tree. One operating point per judge, set by the preregistered rule — the lowest threshold among a judge's benign scores at which its benign false-flag rate is at or below 5% — and three judges excluded for not reaching it. 21 harmful items, so every rate moves in steps of 0.048; no confidence procedure was preregistered and none is reported. The 15 pairs come from 6 judges and are not independent; no multiplicity control was preregistered. Supersedes nothing: E7 on the seven frontier judges is VOID on a defective preregistration that stated its threshold rule in two contradictory directions, those judges are spent, and experiments/e7/RESULT.md records that rather than repairing it. No model was loaded and no API called, so the host refusal at experiments/e2/run/adapters.py:48 is not engaged. Nothing here transfers to E2's guards, to any deployed stack, or to any vendor product.

03Forbidden rescues

Repairs declared unavailable in advance; using one after a failure would breach the recorded commitment.

04Non-claims — what this does not license
05Binding and freshness
Binding
observations.jsonl
The support is the measurement itself: 15 committed pair rows. The commit order is load-bearing — PREREG.md's commit f646e136 precedes the results commit, which is the evidence that the nine hold-out score files were unread when the five predictions were fixed. The upstream score files are pinned by commit and cached under freeze/cache/; they are MIT-licensed and their redistribution is permitted, unlike E6's kernel bytes.
Reviewed
2026-09-10 · window 120 days
Expires
2027-01-08 — after this date the recorded review is overdue; this does not make the claim false
Triggers
  • executable fires when a committed pair row changes without a registry re-review
  • executable fires when the recorded run result or a prediction verdict changes
  • executable fires on any edit to the preregistration — the one file whose contents must not move after the hold-out was read
  • executable fires when the run report changes, including any edit to a stated non-claim
  • executable correction disposition and preserved evidence must remain bound to this review (3 bound sources)
  • manual a pool with intermediate marginals and more than 21 harmful items becomes available, or the upstream repository's scores change at a new commit
Dimensions
visibilityPublicprovenanceMachine-generated, owner-executedsupport roleExecuted outputmaturityExperimental