Recomputing from the committed rows yields any registered quantity different beyond 5e-5, or the pinned kernel bytes are shown not to reproduce all twelve printed Table 3 numbers to the printed decimal, or GDP is shown not to be monotone in one of the five parameters over its published interquartile range, or the measured joint median of 8.6 is shown to lie outside the coupling range this claim reports, or the extraction is shown to have altered the kernel's arithmetic rather than only re-instantiating its own factory source.
Consequence REJECTThis is the condition and consequence recorded in the registry. The vocabulary this is published in defines what each consequence commits the author to.
Historical run only. Formal disposition: experiments/e6/CORRECTION-2026-09-10.md. Frozen inputs, estimator and outputs are unchanged. The expected block below records original run values and must not be used as current confirmatory evidence. Original scope, including interpretations now rejected with this claim: Exactly the 243 rows committed at experiments/e6/results/observations.jsonl, produced 2026-09-10 from one set of kernel bytes pinned by digest in experiments/e6/freeze/sources.json. The five marginals are the published Table 2 quantiles and nothing else; each is discretised to three atoms at those quantiles with weights 0.25/0.50/0.25, a construction chosen after the outcomes were visible and declared as such, which makes the coupling range an inner bound — a finer discretisation of the same marginals can only widen it. scripts/verify_e6.py re-derives all 34 registered quantities from the committed rows alone, with no network and no re-run of the model. This is a re-derivation of public artifacts and not a preregistered test; no prediction here held, because none was made. It transfers no number to any safety classifier, pool, or operating point, and asserts no error in the source paper.
Repairs declared unavailable in advance; using one after a failure would breach the recorded commitment.
- do not re-pin to a different kernel build after seeing an outcome and report the new numbers as this run
- do not change the discretisation, its weights, or the hold-out values after the fact and present the resulting range as this one
- do not describe E6 as preregistered, or any figure in it as a prediction that held
- do not restate the unresolved site-copy estimand as though the public record settled which summary it names
- do not carry any number here across to a guardrail, a classifier, or E2
- Current disposition is REJECT; the historical non-claims below do not restore the rejected assertion. See /corrections/#e6-marginals.
- no error in the source paper is claimed; Table 4 does the joint-preserving computation and does it correctly, and the kernel agrees with the paper wherever both speak
- not a claim that the site's "GDP is 10% higher" figure was computed by composing marginals; two different summaries of the same survey both round to it, the public record does not say which, and what is recorded is that the estimand is not identified from the artifact
- the coupling range is an inner bound under a declared discretisation, not a sharp Fréchet bound; in five dimensions the comonotone corner is attainable but the lower envelope is not a copula, and that optimisation was not solved
- not evidence about the US economy, about AI's economic effects, or about whether any scenario is likely; it is a statement about what a model's inputs determine
- not a guardrail measurement and not transferable to one; it is the same identification structure measured on a different object
- says nothing about the competence or intent of the paper's authors or reviewers
- not a claim that the model composes its five inputs wrongly: the elicited quantities are conditional — the explorer's adoption dial asks what share of the tasks AI *can* do people will use it for — so their product within one respondent is the chain rule and is exact. The coupling this claim varies is the joint distribution across the 10,980 respondents, which no chain rule fixes
- the 243 rows are evaluations of someone else's model at published parameter vectors, not per-item measurements of a guard on an item; they are counted by the evidence ledger as rows this repository produced, and the ledger's single observation-row total must not be read as growth in guard-item measurement. E3 and E3B's 4,800 rows are that kind; these are not