Recomputing from the committed rows yields any registered quantity different beyond 1e-9, or the preregistration's commit is shown not to precede the results commit, or any of the nine hold-out score files is shown to have been retrieved or inspected before f646e136, or the operating-point rule applied in experiments/e7b/run/measure.py is shown to differ from the one PREREG.md states, or an observed joint miss is shown to lie outside the Fréchet interval its own marginals allow.
Consequence REJECTThis is the condition and consequence recorded in the registry. The vocabulary this is published in defines what each consequence commits the author to.
Historical run only. Formal disposition: experiments/e7b/CORRECTION-2026-09-10.md. Frozen inputs, estimator and outputs are unchanged. The expected block below records original run values and must not be used as current confirmatory evidence. Original scope, including interpretations now rejected with this claim: Exactly the 15 rows committed at experiments/e7b/results/observations.jsonl, computed 2026-09-10 from nine published score files in shawnray-research/certified-agent-guardrails at commit 79097583be7786976ea1b9ae79f3ff900d9e66b7 (MIT). Labels are the 35 strings of _inj_goals.json in that tree. One operating point per judge, set by the preregistered rule — the lowest threshold among a judge's benign scores at which its benign false-flag rate is at or below 5% — and three judges excluded for not reaching it. 21 harmful items, so every rate moves in steps of 0.048; no confidence procedure was preregistered and none is reported. The 15 pairs come from 6 judges and are not independent; no multiplicity control was preregistered. Supersedes nothing: E7 on the seven frontier judges is VOID on a defective preregistration that stated its threshold rule in two contradictory directions, those judges are spent, and experiments/e7/RESULT.md records that rather than repairing it. No model was loaded and no API called, so the host refusal at experiments/e2/run/adapters.py:48 is not engaged. Nothing here transfers to E2's guards, to any deployed stack, or to any vendor product.
Repairs declared unavailable in advance; using one after a failure would breach the recorded commitment.
- do not re-threshold, re-calibrate, change the 5% budget, or change the direction of the operating-point rule and report the new numbers as this run
- do not lower the six-pair power floor, or raise it, after the fact
- do not pool the sixteen read judges of E7 and its exploratory record into this panel and report the combined figure as this result
- do not present this as a repair, correction or replacement of the void E7, and do not present E7's frontier judges as still available for a preregistered claim
- do not restate a prediction, narrow it, or drop it from the record
- do not carry these numbers to E2's guards, to a deployed stack, or to a vendor
- Current disposition is REJECT; the historical non-claims below do not restore the rejected assertion. See /corrections/#e7b-pools.
- the Fréchet bound applied here is not this repository's: arXiv:2607.22868v1 states it for an any-flag gate as a proposition, and that artifact's own ensemble_robustness.py reports the vulnerability as correlated across judges. That author found the bound, released the per-item scores that make this checkable, and named the correlation; E7B measures it against a preregistration and claims priority for neither
- 21 harmful items is a small pool; every rate moves in steps of 0.048, no confidence procedure was preregistered, and no point estimate here implies an interval
- 15 pairs drawn from 6 judges are not 15 independent observations, and no multiplicity control was preregistered or is claimed
- one benchmark, one pool, one operating point per judge; AgentDojo injection goals are not a deployed threat model and these are research judges, not a shipped stack
- the three excluded judges were excluded by a pre-stated rule, and their absence is not evidence about them
- not evidence that any specific product, vendor or deployed guardrail stack has correlated failures
- five predictions holding is not a general law: it is one preregistered result on one pool, and the interpretation it supports is that independence understated the joint here, not that it always will