What shipped
The record became an observatory: a thirteen-claim registry (schema v0.3) rendered three ways (ledger, capsule field, six modules under one grammar), an interactive feasible-worlds instrument whose geometry is asserted in CI to 1e-9, clean-clone reproduction of bounds and endpoint witnesses, and a new gate that clones every bound repository to prove each pinned commit is reachable from its default branch — because one binding was found stranded in rewritten history, and byte-watching triggers were structurally blind to it. An independent hostile review filed twelve findings; all twelve were fixed before shipping, one of them a factually false sentence in the module that lectures about expiry. Log 001 carries the lineage of this failure class.
The rating — owner-judged, adversarially informed, no anesthesia
Scores are the owner's judgments after the hostile review, not measurements. Each row names why the score is not higher, which is the only interesting column.
Instrument-grade scaffolding around a research program that has not yet produced its first empirical result. The scaffolding is genuinely rare. It is still scaffolding, and thirty days from now the only question that matters is whether something real got lifted.
What would make this instrument dangerous
Dangerous in the only sense this record endorses: menace pointed at claims, never at people. The site's machinery — envelopes, bounds, witnesses, expiry — currently prosecutes only its own claims. The move that turns it from a portfolio into a threat to overclaiming is to point it outward: take published AI-safety numbers from the wild, compute what their stated evidence actually licenses under the Fréchet bounds, and publish the interval next to the headline number with witnesses at both ends — a claim autopsy, powered by exactly the E2-style measurement this program has frozen but not yet run. The standard, made visible, applied where it stings. That requires E2 to exist first, which is why the lookback below has teeth.
On or after that date, answer each question against evidence that exists outside this page. Anyone can run this checklist; none of it requires trusting the owner.
- Did the standing gate stand? The weekly verify runs (five expected by then) are public in the Actions history. Any red run, and what fired.
- Did E2 move? Either a conforming observation set exists, or a documented failure-to-obtain exists. If neither, the "next artifact" promise on /now/ is thirty days stale and this entry says so.
- Were the open QA items closed? Keyboard-only and screen-reader passes were disclosed as not-run at ship time. After thirty days, an open disclosure becomes a broken promise; the accessibility score above drops to 5 until they run.
- Did any external falsifier arrive? Corrections, issues, or mail against any claim — and what the registry did about them.
- Do any decay clocks show amber? None should before December; amber earlier means a review was triggered, and the ledger must show the re-review.
- Did the claim count move for the right reason? New claims should arrive because evidence arrived — not because surfaces multiplied.
Pre-registered conclusion if nothing moved: the observatory spent its first month documenting stasis beautifully — an instrument without an experiment — and the research-threat score falls from 5 to 3. That sentence is written now, while it can still be falsified by work.