The first rows this record produced, and what they could not answer
On 2026-09-06 two pilots ran on hardware I own and produced 4,800 per-item
observation rows — the first measurements this repository has made rather than
recounted. Both failed their primary pre-registered prediction, and both are registered
with that failure in the commitment: E3-001 and
E3B-001. E3 predicted an excess joint miss whose interval
excluded zero and returned +0.0018 with a 95% interval of [−0.00096, +0.00706]. E3B was
pre-registered to widen the identified set past ten points and returned a set of width
zero. scripts/verify_e3.py re-derives every registered number, including the
bootstrap interval, from the committed rows alone on every push.
Why neither could answer its question is the useful part. In E3 both guards missed 96–98% of the pool, so the two marginals alone pinned the joint miss to a 1.75-point interval. In E3B one guard missed nothing, so the marginals fixed the joint miss exactly and the interval had width 0. Two runs, two degenerate identified sets, at opposite ends of the marginal range. That the identified set is wide in the middle and degenerate at the extremes — and that a roughly forty-item pre-scoring check would have stopped both runs — is a hypothesis these two pilots suggest and neither tested. It is registered as a non-claim in those words, not as a finding. E3B is a new experiment with its own freeze, not a re-run or a correction of E3; both results stand.
Trial IV is live and no human has answered it
A twenty-minute exercise, PC-001, went live at /trials/necromancer/ on 2026-09-03. It asks whether a person can be made better at refusing to rescue their own claim once the evidence turns against it. The instrument is frozen: the item text, the scoring key digest, the two-arm assignment for ten enrolment slots, the primary outcome and the continue / narrow / kill thresholds were all fixed on 2026-09-02, before any human saw the page, and the seven files they cover are hash-pinned. The scoring script refuses to conclude anything below eight receipts.
The honest state is that zero people have taken it. The invitation has not been sent, so no enrolment exists, no receipt exists, and the pilot has produced no observation row. Nothing on this page or in the repository can change that: sending the invitation is a human action and it has not happened. Until it does, PC-001 is a hypothesis with a frozen instrument and no data, and it is recorded here as exactly that rather than as progress. This entry does not claim the exercise works, that it teaches anything, or that its question is novel; it claims only that the instrument exists, is frozen, and is unanswered.
The Missing Column Census — the field question, made falsifiable
The question this record keeps asking is narrow: which public guardrail evaluations preserve joint evidence for a stated static composition, rather than only per-system marginals? In a bounded, single-reviewer inventory, twenty artifacts have now been examined against primary sources. The census retains its 2026-08-27 cutoff and was corrected on 2026-08-30 after review found a qualifying artifact with a pre-cutoff source commit. Fourteen document a shared item set and a common event definition; the stricter ladder is 14 / 12 / 0 for shared basis / no stated threshold mismatch / documented matched thresholds with full exposure. Five preserve heterogeneous joint-evidence artifacts: four print a composition result and two release aligned per-item outcomes from which one is directly computable; one artifact does both. The prior 19 / 13 / 4 envelope is recorded as rejected rather than silently rewritten. A narrower commercial-API clause was withdrawn after fresh-context adversarial review; the record preserves that correction rather than rescuing the claim. The census also makes its named-products sensitivity visible (19 examined / 13 shared-basis / 5 joint-evidence artifacts), rather than hiding an ambiguity in the locked wording. The counts are recomputed from the census file in CI, so the headline is exactly as falsifiable as the rest of this record.
The E2 instrument ran, and reported a null
The full E2 measurement path — schema-conforming rows, the shared-item rule, one frozen reduction, the excess-joint-failure estimand, and the pre-registered negative controls — was exercised against three toy mechanisms. Sixty-six rows passed the real E2 validator with zero violations; the observed dependence was indistinguishable from independence at twenty-two items, and the controls behaved exactly. It is not E2, and it says so on its face; what moved is that the instrument is proven and the cost of the real study is now a number.
E2 — the shared-item measurement contract
CC-Framework's first empirical rung was frozen before any dataset was inspected, and no conforming dataset has been collected. On 2026-08-23 the measurement pipeline was rehearsed end to end on three synthetic mechanisms — a documented dry run, not E2 — which validated the instrument and put a number on what real E2 costs: about 1,097 shared items run through every guardrail to estimate the rates to within five points. That is the bar the next artifact must clear. E2 stays frozen and untested until it does.
Incidence, not existence — the receipt-kernel gap
Ghost-Ark's collision claim stands in contracted form: possible, not observed (a sample of 64 real payloads carried zero pathology classes). The gap its thesis names after that result: incidence, not existence — does any real consumer population currently depend on a distinction some deployed canonicalizer destroys? (Separately, its falsifier ledger keeps F3 — that the consumer set is stable in practice — open and, in its own words, the most under-attacked of the five.)
The observatory redesign — v0.4
This site became the Evidence Observatory: a six-module question grammar, the claim
observatory view, an interactive feasible-worlds instrument, and geometry assertions that
make a figure that lies about probability fail the build. The 2026-08-23 release
contained 13 claims. That is a dated snapshot, not the current count: the
ledger renders the current record directly from
claims.yaml. One stale binding was found and re-pinned in the process; the
log has the details.
Module 004 — the attestation boundary
What can enclave attestation prove about a file's life before the enclave first touched it? The standing conjecture: nothing, by itself. The page exists, says PLANNED on its face, and carries no evidence because there is none to carry.
This page is hand-written and owner-attested; it carries no evidence markers and borrows none. When it goes stale, it is replaced, not accreted — the previous state survives in git history, which is where history belongs.