Pranav Bhave

Now · updated 2026-09-08 · owner-attested, dated, replaced when stale

Can dependence evidence be measured at feasible cost — before it has to be assumed?

One claim should become more inspectable every working day. This page names the claim currently on the bench, and the honest state of everything around it.

Measured · both predictions failed · 2026-09-06

The first rows this record produced, and what they could not answer

On 2026-09-06 two pilots ran on hardware I own and produced 4,800 per-item observation rows — the first measurements this repository has made rather than recounted. Both failed their primary pre-registered prediction, and both are registered with that failure in the commitment: E3-001 and E3B-001. E3 predicted an excess joint miss whose interval excluded zero and returned +0.0018 with a 95% interval of [−0.00096, +0.00706]. E3B was pre-registered to widen the identified set past ten points and returned a set of width zero. scripts/verify_e3.py re-derives every registered number, including the bootstrap interval, from the committed rows alone on every push.

Why neither could answer its question is the useful part. In E3 both guards missed 96–98% of the pool, so the two marginals alone pinned the joint miss to a 1.75-point interval. In E3B one guard missed nothing, so the marginals fixed the joint miss exactly and the interval had width 0. Two runs, two degenerate identified sets, at opposite ends of the marginal range. That the identified set is wide in the middle and degenerate at the extremes — and that a roughly forty-item pre-scoring check would have stopped both runs — is a hypothesis these two pilots suggest and neither tested. It is registered as a non-claim in those words, not as a finding. E3B is a new experiment with its own freeze, not a re-run or a correction of E3; both results stand.

Hypothesis · open · 2026-09-04

Trial IV is live and no human has answered it

A twenty-minute exercise, PC-001, went live at /trials/necromancer/ on 2026-09-03. It asks whether a person can be made better at refusing to rescue their own claim once the evidence turns against it. The instrument is frozen: the item text, the scoring key digest, the two-arm assignment for ten enrolment slots, the primary outcome and the continue / narrow / kill thresholds were all fixed on 2026-09-02, before any human saw the page, and the seven files they cover are hash-pinned. The scoring script refuses to conclude anything below eight receipts.

The honest state is that zero people have taken it. The invitation has not been sent, so no enrolment exists, no receipt exists, and the pilot has produced no observation row. Nothing on this page or in the repository can change that: sending the invitation is a human action and it has not happened. Until it does, PC-001 is a hypothesis with a frozen instrument and no data, and it is recorded here as exactly that rather than as progress. This entry does not claim the exercise works, that it teaches anything, or that its question is novel; it claims only that the instrument exists, is frozen, and is unanswered.

Corrected · 2026-08-30

The Missing Column Census — the field question, made falsifiable

The question this record keeps asking is narrow: which public guardrail evaluations preserve joint evidence for a stated static composition, rather than only per-system marginals? In a bounded, single-reviewer inventory, twenty artifacts have now been examined against primary sources. The census retains its 2026-08-27 cutoff and was corrected on 2026-08-30 after review found a qualifying artifact with a pre-cutoff source commit. Fourteen document a shared item set and a common event definition; the stricter ladder is 14 / 12 / 0 for shared basis / no stated threshold mismatch / documented matched thresholds with full exposure. Five preserve heterogeneous joint-evidence artifacts: four print a composition result and two release aligned per-item outcomes from which one is directly computable; one artifact does both. The prior 19 / 13 / 4 envelope is recorded as rejected rather than silently rewritten. A narrower commercial-API clause was withdrawn after fresh-context adversarial review; the record preserves that correction rather than rescuing the claim. The census also makes its named-products sensitivity visible (19 examined / 13 shared-basis / 5 joint-evidence artifacts), rather than hiding an ambiguity in the locked wording. The counts are recomputed from the census file in CI, so the headline is exactly as falsifiable as the rest of this record.

Moved · 2026-08-23

The E2 instrument ran, and reported a null

The full E2 measurement path — schema-conforming rows, the shared-item rule, one frozen reduction, the excess-joint-failure estimand, and the pre-registered negative controls — was exercised against three toy mechanisms. Sixty-six rows passed the real E2 validator with zero violations; the observed dependence was indistinguishable from independence at twenty-two items, and the controls behaved exactly. It is not E2, and it says so on its face; what moved is that the instrument is proven and the cost of the real study is now a number.

Frozen · untested — by design

E2 — the shared-item measurement contract

CC-Framework's first empirical rung was frozen before any dataset was inspected, and no conforming dataset has been collected. On 2026-08-23 the measurement pipeline was rehearsed end to end on three synthetic mechanisms — a documented dry run, not E2 — which validated the instrument and put a number on what real E2 costs: about 1,097 shared items run through every guardrail to estimate the rates to within five points. That is the bar the next artifact must clear. E2 stays frozen and untested until it does.

Open · under-attacked

Incidence, not existence — the receipt-kernel gap

Ghost-Ark's collision claim stands in contracted form: possible, not observed (a sample of 64 real payloads carried zero pathology classes). The gap its thesis names after that result: incidence, not existence — does any real consumer population currently depend on a distinction some deployed canonicalizer destroys? (Separately, its falsifier ledger keeps F3 — that the consumer set is stable in practice — open and, in its own words, the most under-attacked of the five.)

Landed · 2026-08-23

The observatory redesign — v0.4

This site became the Evidence Observatory: a six-module question grammar, the claim observatory view, an interactive feasible-worlds instrument, and geometry assertions that make a figure that lies about probability fail the build. The 2026-08-23 release contained 13 claims. That is a dated snapshot, not the current count: the ledger renders the current record directly from claims.yaml. One stale binding was found and re-pinned in the process; the log has the details.

Planned · question stated

Module 004 — the attestation boundary

What can enclave attestation prove about a file's life before the enclave first touched it? The standing conjecture: nothing, by itself. The page exists, says PLANNED on its face, and carries no evidence because there is none to carry.

This page is hand-written and owner-attested; it carries no evidence markers and borrows none. When it goes stale, it is replaced, not accreted — the previous state survives in git history, which is where history belongs.