Pranav Bhave
Observatory — the registry as one field · generated from claims.yaml Schema v0.4 · 21 claims · 2 contradicted · 18 supported within scope · 1 publicly untested

Claim observatory

Every claim in this record, as a capsule: its proposition, its evidence binding, its decay clock, the boundary it will not cross, and the condition that would defeat it. A solid cyan rail marks a public evidence binding; capsules without one carry a dashed neutral rail and an attested chip that says so; the clock turns amber when a review window is three-quarters spent and red when it lapses. There is no aggregate score here and never will be — a quiet trigger is quiet, not "healthy", and these envelopes do not average into one number.

CC-001

Supported within scope

cc.kernel.strict.frechet_bounds computes Fréchet–Hoeffding endpoint bounds for composed binary guardrail failure from declared marginals; for marginals (0.10, 0.10) it returns [0.0, 0.10] for the AND event and [0.10, 0.20] for OR.

Machine-generated, owner-executed · Released reviewed 2026-08-24 · window 120d falsifier consequence: REJECT · 2 forbidden rescues executablemanual

CC-003

Supported within scope

CC-Framework's E1 dependence-evidence study established constructively that distributions with identical singleton failure rates (0.5) and identical pairwise overlaps (0.25) can differ in three-way failure probability — 0 under even parity, 0.25 under odd parity — so measuring every pair does not identify a three-guardrail stack. E1 is synthetic; its recorded decision is Narrow.

Owner-authored, public · Released reviewed 2026-08-24 · window 120d falsifier consequence: REJECT · 2 forbidden rescues executablemanual

CC-004

Supported within scope

For marginals (0.10, 0.10) and the both-fail query, cc.kernel.strict identify() returns the sharp interval [0.0, 0.10] together with endpoint witness distributions that sum to one, are nonnegative, satisfy both marginal constraints, and attain each endpoint — verified to 1e-9 from a clean clone.

Machine-generated, owner-executed · Released reviewed 2026-08-24 · window 120d falsifier consequence: REJECT · 2 forbidden rescues executablemanual

CC-005

Supported within scope

CC-Framework's E2 — the shared-item empirical guardrail pilot — is governed by a measurement contract frozen before any dataset was inspected; no conforming dataset has been collected, so E2 is untested, and the contract's conformance checks are executable.

Owner-authored, public · Released reviewed 2026-08-24 · window 120d falsifier consequence: NARROW · 3 forbidden rescues executablemanual

CC-006

Supported within scope

CC-Framework's E2 measurement pipeline was rehearsed end to end on three synthetic guardrail mechanisms: 66 observation rows conform to the frozen E2 schema with zero validator violations, the pre-registered negative controls behave (the common-cause control attains pA(1-pA) to 1e-9), and the finding records the result as a null at n=22 with a quantified cost for real E2 (~1097 shared items for epsilon=0.05). E2 itself is not run.

Owner-authored, public · In development reviewed 2026-09-07 · window 120d falsifier consequence: REJECT · 2 forbidden rescues executablemanual

REL-001

Publicly untested

HYPOTHESIS — Among English-reading professionals who made or materially prepared a go/hold recommendation for an AI-based system in paid work during the prior 12 months, an E2-derived dependence-aware uncertainty presentation will increase HOLD decisions by at least 15 percentage points relative to a stipulated-independence presentation of the same evidence.

attested — no public artifact — stated on the owner's responsibility
Owner-attested · Experimental reviewed 2026-08-24 · window 30d falsifier consequence: REJECT · 4 forbidden rescues manual

MC-001

Supported within scope

As of 2026-08-27, among 20 public guardrail evaluations meeting the Missing Column Census's frozen inclusion criteria (v1), 14 establish a shared item set and a common event definition; 5 provide one of the census's declared joint-evidence artifacts: a printed composition result or per-item outcomes that make one directly computable.

Owner-authored, public · Experimental reviewed 2026-09-14 · window 60d falsifier consequence: REJECT · 4 forbidden rescues executablemanual

MC-002

Supported within scope

On the 170-prompt per-item subset publicly released with BELLS's 2025 misuse-detection evaluation (pinned commit below), the five specialized supervisors — Lakera Guard, Prompt Guard, LangKit, NeMo Guardrails, LLM Guard — have an OR-union of released binary verdict columns that flags 73 of the 82 prompts labelled harmful, leaving 9 all-miss (11.0%). The product of their individual miss rates is 3.5%; the release-recomputed all-miss rate is about 3.1× that independence plug-in on this file. The same OR-union flags 19 of the 50 benign prompts.

Machine-generated, owner-executed · Experimental reviewed 2026-08-28 · window 60d falsifier consequence: REJECT · 4 forbidden rescues executablemanual

MC-003

Supported within scope

For k guardrails scored on a common item set at a common operating point under block-on-any composition, the rate at which every guard misses the same item is fixed by the published per-guard miss rates only up to the interval [max(0, Σp − (k−1)), min p], and every point of that interval is attained by some joint law with those marginals. Marginals prove non-degradation of a static OR composition relative to its best member, but, when that identified set is non-degenerate, they do not identify strictly positive incremental benefit. They also identify a positive lower bound on static benign union-flag rate when any member has a positive benign flag rate. On the BELLS 2025 partial per-item release — 82 prompts labelled harmful — the exact finite identified set is {0/82 … 12/82}; MC-002 recomputes 9/82 all-miss from the hash-verified released verdict file. Marginal catch counts plus the union identify only bounds on each guard's exclusive full-stack coverage; the registered leave-one-out unions identify the realized values: NeMo 18, Lakera 3, and LangKit, Prompt Guard, and LLM Guard 0 on this stratum.

Machine-generated, owner-executed · Experimental reviewed 2026-08-30 · window 60d falsifier consequence: REJECT · 5 forbidden rescues executablemanual

MC-004

Supported within scope

On the per-item guard-verdict files publicly released with Multimodal Safeguard Bench's full_run (pinned commit below; six verdict files and the release's two printed-metric files hash-verified), the three guards — Llama Guard 4, Llama Guard 3 Vision, ShieldGemma 2 — have harness-normalized native `unsafe` labels whose static Boolean OR is 1 on 192 of 200 harmful-labelled text items (8 all-zero rows, 4.0%) and all 200 harmful-labelled image items (0 all-zero rows), while the same OR is 1 on 32 of 250 benign-labelled text items (12.8%) and every one of 250 benign-labelled image items. Llama Guard 3 Vision alone has a 1-bit on all 250 released benign image items. The complete 2^3 native-label pattern table, leave-one-out bit ORs, and unique bit contributions for each released stratum are bound below. The release prints per-guard metrics and two two-guard compositions; it prints no three-guard bit OR, no all-zero-row count, no leave-one-out table, and no pattern decomposition — and every recomputed quantity that overlaps what it does print is asserted equal to the printed value.

Machine-generated, owner-executed · Experimental reviewed 2026-08-31 · window 60d falsifier consequence: REJECT · 6 forbidden rescues executableexecutableexecutableexecutableexecutableexecutableexecutableexecutableexecutableexecutableexecutablemanual

AF-001

Supported within scope

Anthropic's AI Fluency Index reports overall prevalences for the 11 of its framework's 24 behaviours that are directly observable in conversations, and percentage-point differences for conversations that produced an artifact, but not the artifact share and not which comparison group the differences are against. Those published numbers therefore do not identify the subgroup rates — but they do bound them. Under every feasible artifact share and under either reading of the comparison, fact-checking is between 30% and 43% less prevalent in artifact conversations, a far larger effect than the stated 3.7 percentage points suggests on a base rate of 8.7%. The three behaviours that rise in artifact conversations are Description behaviours; the three that fall are Discernment behaviours.

Machine-generated, owner-executed · Experimental reviewed 2026-08-28 · window 90d falsifier consequence: NARROW · 4 forbidden rescues executablemanual

GA-001

Supported within scope

Ghost-Ark is a verifier and measurement harness for the provenance limits of AI-governance receipts — a research artifact of the S2 Lab, Penn State — whose stated thesis is that receipt soundness is a ternary relation Sound(C, Σ, P): a receipt identifies an execution only up to the kernel of its whole parse → canonicalize → digest pipeline, so soundness does not persist by default as the pathology alphabet grows or the consumer set widens.

Lab repository, public · In development reviewed 2026-08-24 · window 120d falsifier consequence: NARROW · 2 forbidden rescues executable

GV-001

Supported within scope

Ghost Visualizer is a seven-scene React/Vite visual essay ("Why AI Safety Scores Lie") that computes marginals, OR composition, Fréchet–Hoeffding bounds, endpoint witnesses, and a receipt object from deterministic sample rows, with an Evidence Mode for inspection.

Owner-authored, public · In development reviewed 2026-08-24 · window 120d falsifier consequence: NARROW · 2 forbidden rescues executablemanual

GCE-001

Supported within scope

GCE — the Guardrail Composability Explorer — is a coursework MVP demo (AI-285) for toggling guardrails and observing composed behavior. Its front-door composability-coefficient framing is superseded by CC-Framework's dependence-aware partial identification, whose metric taxonomy classifies the older coefficient family as legacy/deprecated. GCE is preserved as intellectual lineage, not current theory.

Owner-authored, public · Superseded reviewed 2026-08-24 · window 365d falsifier consequence: NARROW · 2 forbidden rescues executablemanual

SITE-001

Supported within scope

This site's palette token pairs were computed to pass WCAG AA contrast (most AAA), and manual QA passes covered light/dark themes, desktop and mobile layouts, keyboard focus, reduced motion, and no-JS rendering, as recorded in DESIGN.md.

Owner-verified · Released reviewed 2026-08-31 · window 120d falsifier consequence: NARROW · 2 forbidden rescues executablemanual

SITE-002

Supported within scope

This site's claim registry is enforced in CI: verify_claims.py validates schema and bindings, executes the executable review triggers against live evidence, and fails the build when any claim passes its freshness window — on every push and weekly.

Owner-verified · Released reviewed 2026-08-30 · window 120d falsifier consequence: REJECT · 3 forbidden rescues executablemanual

E3-001

Supported within scope

E3 — the first pilot this repository ran itself — scored two ungated classifiers on 400 harmful and 800 benign items and produced 2,400 committed observation rows. Its primary pre-registered prediction FAILED: excess joint miss was +0.0018 with a 95% bootstrap CI of [-0.00096, +0.00706], which includes zero. Both guards missed almost everything on this pool (0.9825 and 0.9625), so the Fréchet interval the two marginals allow is [0.9450, 0.9625] — 1.75 percentage points wide — and the observed joint miss of 0.9475 lies inside it. The prediction that it would lie inside HELD; the difficulty-stratification prediction was NOT COMPUTED and remains open.

Machine-generated, owner-executed · Experimental reviewed 2026-09-08 · window 120d falsifier consequence: REJECT · 4 forbidden rescues executableexecutableexecutablemanual

E3B-001

Supported within scope

E3B put the same two classifiers on the attack family they were built for — 400 real prompt injections — and produced a further 2,400 committed observation rows. Its primary prediction FAILED and so did the prediction the redraw existed to test. One guard missed nothing (0.0000, catching 400 of 400) and the other missed 0.3975, so the Fréchet interval the marginals allow is [0.0000, 0.0000] — zero points wide — the observed joint miss is exactly 0.0000, and the bootstrap CI is the degenerate [0, 0]. E3B was pre-registered to produce an interval wider than 10 percentage points; it produced one narrower than E3's. The prediction that the observed joint miss would lie inside the interval HELD, trivially.

Machine-generated, owner-executed · Experimental reviewed 2026-09-08 · window 120d falsifier consequence: REJECT · 4 forbidden rescues executableexecutableexecutablemanual

E6-001

Contradicted

REJECTED AS STATED on 2026-09-10. The permutation routine moves unequal-mass atoms. Only 16 of 1,296 constructions preserve the specified marginal weights. The advertised 2.12–18.07 fixed-marginal range is unsupported and is rejected. No replacement range is asserted. Historical assertion retained for audit, not an accepted current finding: E6 extracted the client-side model kernel that anthropic.com/institute/econ-scenarios ships (Turbopack module 303412, sha256 0a75763f…), gated it against the twelve printed numbers of Table 3 of the Anthropic Institute's Working Paper 2026-02, and evaluated it only at parameter vectors built from the five marginal quantiles that paper prints in Table 2. The kernel reproduces all twelve Table 3 numbers to the printed decimal, and GDP is strictly increasing in each of the five parameters over their published interquartile ranges. Holding the four parameters the note to Table 4 names at that note's values, the same five marginals admit a median 2030 GDP anywhere in [2.12, 18.07] percent above the no-AI path across the 1,296 rank-permutation couplings of their published quartiles: 9.88 at the comonotone corner, 8.32 under independence, against the 8.6 the paper obtained by running each of 3,259 respondents' own five-vector. The marginals fix an interval fifteen points wide, not a point, and the measured joint lies strictly inside it.

Machine-generated, owner-executed · Experimental reviewed 2026-09-10 · window 120d falsifier consequence: REJECT · 5 forbidden rescues executableexecutableexecutableexecutableexecutableexecutableexecutablemanual

E7B-001

Contradicted

REJECTED AS STATED on 2026-09-10. Calibration uses the all-nine shared pool instead of each judge’s own pool, changing one threshold from 0.95 to 1.00. Pair evaluation also uses the global intersection. The preregistration-conforming confirmation claim is rejected; the original HELD labels are historical outputs. Historical assertion retained for audit, not an accepted current finding: E7B is the first non-degenerate preregistered measurement this repository has produced, and all five of its predictions held on data unread when they were fixed. Against a preregistration frozen at f646e136 before any hold-out byte was retrieved, six of nine published guardrail-judge score sets reached a 5% benign false-flag budget on their shared 96-item AgentDojo pool (21 injection goals, 75 benign) and yielded 15 non-degenerate pairs. Median excess joint miss was +0.2018 and mean +0.1852 (P1, P2 HELD); the median pair's observed joint miss sat 0.900 of the way to the Fréchet upper bound its own marginals allow (P3 HELD); every observed joint lay inside its interval (P4 HELD); and every one of the 15 pairs had an observed joint miss above the independence product (P5 HELD). On the modal pair two judges each miss about half the injection goals, independence predicts a 0.249 both-miss rate, and the observed rate is 0.476 — exactly the upper bound, meaning one judge's misses are a subset of the other's. Seven of the fifteen pairs sit at that bound.

Machine-generated, owner-executed · Experimental reviewed 2026-09-10 · window 120d falsifier consequence: REJECT · 6 forbidden rescues executableexecutableexecutableexecutableexecutableexecutableexecutablemanual

The non-claims wall

The outer boundary of the record: everything these claims refuse to support, collected in one place. A claim without a stated non-claim is a claim that has not found its edge yet.

CC-001
  • does not certify any stacked system as safe
  • does not recover or estimate the unknown dependence
  • does not select a point inside the returned interval
CC-002
  • the manifest's existence does not validate the claims it scopes
CC-003
  • the parity construction demonstrates possibility, not frequency, in real guardrail stacks — no claim of empirical prevalence
  • nothing here validates the framework's practical or deployed value; "we validated CC" is prohibited at every rung of the repository's evidence ladder
  • the product-baseline failure directions observed in E1 are properties of the tested generators, not universal constants
CC-004
  • endpoint witnesses are feasible mathematical worlds, not observed systems
  • attainment does not select a point inside the interval or estimate the true dependence
CC-005
  • freezing a contract establishes discipline, not results; E2 remains untested until conforming observations exist
  • conformance is not a clean bill of health, and the contract says so
  • "we validated CC" is prohibited at every rung of the evidence ladder
CC-006
  • this is not E2 and does not move E2 off untested; the mechanisms are toy filters in the repository, not deployed guardrails
  • the observed dependence is a null at this scale, not evidence of independence or of any real-guardrail dependence
  • validating the instrument establishes nothing about the safety, representativeness, or deployment behavior of any real system
REL-001
  • no practitioner decision, E2-derived participant packet, consent record, recruitment, or study outcome exists yet
  • this does not establish that dependence-aware evidence changes real deployment decisions, is generally useful, or justifies its cost
  • this is not a customer-demand, payment, adoption, safety, or certification claim
MC-001
  • does not claim any evaluated guardrail stack performs poorly; an unfilled column is a reporting fact, not a performance finding
  • does not claim the unmeasured joint statistics would reveal dependence; measuring instead of assuming is the point
  • does not audit the quality of any per-system evaluation beyond the fields each row records
  • does not claim exhaustive coverage beyond the corrected, examined artifact set or coverage after the stated date
MC-002
  • the subset is author-selected with an unstated rule — nothing here estimates any system's true rate, and no confidence interval is offered because the sampled population is undefined
  • not a ranking or endorsement of any vendor; the released verdicts are at unstated default configurations
  • says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release
  • the release-recomputed-to-plug-in ratio describes this subset's arithmetic, not a general law of guardrail dependence
  • not an observed deployed five-guard stack; interpreting this static OR aggregation as a stack requires the separate full-exposure, parallel, fixed-operating-point assumptions
MC-003
  • no new mathematics is claimed; the bounds are Fréchet's and the lower endpoint is Bonferroni's
  • not a claim that guardrails in general fail together — this stratum cannot separate shared blind spots from prompt-difficulty heterogeneity
  • not a vendor evaluation; the released verdicts are at unstated default configurations, and one supervisor fires exactly once in the 170 released rows
  • the identified interval is what the marginals leave open, not a prediction about any deployed stack
  • leave-one-out unions identify only exclusive full-stack coverage; they do not identify pairwise or higher-order overlap, Shapley values, or causal contribution
  • says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release
MC-004
  • not a shared-event catch statistic: the common `blocked` bit is a harness normalization of distinct native predicates, and no source-defined translation to a common event E has been identified
  • not a population estimate — the items derive from HarmBench and XSTest under the release's own construction, and no interval is offered because no sampled population is defined
  • not a ranking, endorsement, or indictment of any guard or vendor; the verdicts are at native, unmatched operating rules
  • not an observed deployed three-guard stack; interpreting the OR aggregation as a stack requires the separate full-exposure, parallel, fixed-operating-point assumptions
  • says nothing about the release's carrier-prompt, adversarial-UAP, or cross-VLM runs, which are separate artifacts with their own contracts
  • the counts are about the committed verdict bytes; the release's own changelog documents that ShieldGemma 2's image scores are sensitive to the text-rendering stack, so nothing here predicts what any guard would do under a different rendering environment
  • the zero three-guard image all-zero-bit count is a counting fact about these 200 released image items, not evidence of general image attack safety
AF-001
  • not a criticism of the Index; the observability boundary and the correlational limits are the report's own declared statements
  • not a claim about any individual's competence, and not a measure of anyone's diligence
  • not evidence that artifact production causes less checking — the published marginals cannot separate a behaviour change from a measurement change
  • no access to the underlying conversations; this is arithmetic on a published summary
GA-001
  • a verifying receipt does not establish that the governed action was safe, authorized, or semantically correct
  • kernel collisions in real canonicalizers are demonstrated as possible, not as prevalent — the repository's E12 sample found 0 of 64 real payloads carrying any pathology class
  • not hardened for deployment; not post-quantum secure
GV-001
  • not an AI-safety certificate; no deployment-readiness claim
  • its own public-readiness review scores it for private serious-contact demo use; the site therefore links source rather than exhibiting its media as a flagship
GCE-001
  • supersession is a statement about theoretical framing, not about the correctness of GCE's code or its value as coursework
  • no current-research claim is made by or for GCE
SITE-001
  • no field CLS/LCP/CrUX measurements are claimed
  • "zero third-party runtime requests" describes page code, not hosting infrastructure (GitHub Pages, DNS, TLS remain dependencies)
SITE-002
  • a green run verifies registry consistency and quiet triggers, never the truth of any claim's content
  • an UNDETERMINED run (exit 2) blocks the build because a source was never reached; it is not a finding about any binding, and must not be read or recorded as a failed check
  • triggers watch file content; semantic drift outside watched files remains a manual review event, and the ledger says so
E3-001
  • the 2,400 rows prove the instrument runs end to end; they do not prove it measures what the programme says it measures, and on this pool the identified set was nearly a point, so it measured almost nothing
  • not evidence that either classifier is good or bad; two research models, one pool, one operating point each
  • the null is a null at this scale on this pool, not evidence of independence and not evidence about any real guardrail's dependence
  • says nothing about E2, its three frozen guards, its pools, or its operating points
  • no vendor, product, deployed stack, or population is described
E3B-001
  • a zero-width identified set means the marginals already fixed the joint miss; it is not a measurement that the two guards fail independently
  • G1 catching 400 of 400 is not evidence that it is a good guardrail, and contamination is unverified rather than excluded
  • G2 missing 159 of 400 is not evidence that it is a bad one
  • 2,400 further rows prove the instrument runs; on this pool the question was degenerate, so they measure almost nothing
  • the sharpening these two pilots suggest — that marginal-only reporting is uninformative in the middle of the marginal range and fully informative at its extremes — is a hypothesis and an engineering design gate, not a result of either pilot; neither run tested it
  • no vendor, product, deployed stack, or population is described
E6-001
  • Current disposition is REJECT; the historical non-claims below do not restore the rejected assertion. See /corrections/#e6-marginals.
  • no error in the source paper is claimed; Table 4 does the joint-preserving computation and does it correctly, and the kernel agrees with the paper wherever both speak
  • not a claim that the site's "GDP is 10% higher" figure was computed by composing marginals; two different summaries of the same survey both round to it, the public record does not say which, and what is recorded is that the estimand is not identified from the artifact
  • the coupling range is an inner bound under a declared discretisation, not a sharp Fréchet bound; in five dimensions the comonotone corner is attainable but the lower envelope is not a copula, and that optimisation was not solved
  • not evidence about the US economy, about AI's economic effects, or about whether any scenario is likely; it is a statement about what a model's inputs determine
  • not a guardrail measurement and not transferable to one; it is the same identification structure measured on a different object
  • says nothing about the competence or intent of the paper's authors or reviewers
  • not a claim that the model composes its five inputs wrongly: the elicited quantities are conditional — the explorer's adoption dial asks what share of the tasks AI *can* do people will use it for — so their product within one respondent is the chain rule and is exact. The coupling this claim varies is the joint distribution across the 10,980 respondents, which no chain rule fixes
  • the 243 rows are evaluations of someone else's model at published parameter vectors, not per-item measurements of a guard on an item; they are counted by the evidence ledger as rows this repository produced, and the ledger's single observation-row total must not be read as growth in guard-item measurement. E3 and E3B's 4,800 rows are that kind; these are not
E7B-001
  • Current disposition is REJECT; the historical non-claims below do not restore the rejected assertion. See /corrections/#e7b-pools.
  • the Fréchet bound applied here is not this repository's: arXiv:2607.22868v1 states it for an any-flag gate as a proposition, and that artifact's own ensemble_robustness.py reports the vulnerability as correlated across judges. That author found the bound, released the per-item scores that make this checkable, and named the correlation; E7B measures it against a preregistration and claims priority for neither
  • 21 harmful items is a small pool; every rate moves in steps of 0.048, no confidence procedure was preregistered, and no point estimate here implies an interval
  • 15 pairs drawn from 6 judges are not 15 independent observations, and no multiplicity control was preregistered or is claimed
  • one benchmark, one pool, one operating point per judge; AgentDojo injection goals are not a deployed threat model and these are research judges, not a shipped stack
  • the three excluded judges were excluded by a pre-stated rule, and their absence is not evidence about them
  • not evidence that any specific product, vendor or deployed guardrail stack has correlated failures
  • five predictions holding is not a general law: it is one preregistered result on one pool, and the interpretation it supports is that independence understated the joint here, not that it always will

Falsifiers and forbidden rescues

A falsifier can otherwise be evaded by changing the proposition after observing the result. Each record fixes its defeat condition and consequence in advance, then lists the reinterpretations that cannot keep it standing. An explicit [] means no meaningful post-falsification rescue applies.

CC-001

FalsifierA clean-clone execution of the bound module on marginals (0.10, 0.10) returns AND or OR bounds other than the recorded endpoints, beyond the stated tolerance.

ConsequenceREJECT

Forbidden rescues

  • do not substitute a different input, revision, or numerical tolerance after the failure
  • do not treat a passing wrapper or test harness as evidence for the stated outputs if direct execution disagrees
CC-002

FalsifierThe bound manifest no longer maps its public claims to validation lanes, supporting files, and explicit non-claims as stated.

ConsequenceREJECT

Forbidden rescues

  • do not call partial or undocumented mappings complete manifest coverage
  • do not substitute a different, unbound checklist or README after the manifest fails this condition
CC-003

FalsifierThe frozen E1 construction fails to retain the stated equal singleton and pairwise values while yielding the stated different three-way probabilities (0 under even parity and 0.25 under odd parity), or the bound study does not record the stated Narrow decision.

ConsequenceREJECT

Forbidden rescues

  • do not switch to different marginals, overlaps, or generators after failure and call it the same construction
  • do not turn failure of the concrete construction into a generic possibility claim without a new, stated construction
CC-004

FalsifierBound identify() execution fails any stated interval, witness feasibility constraint, marginal constraint, or endpoint-attainment assertion at the declared tolerance.

ConsequenceREJECT

Forbidden rescues

  • do not replace a failed endpoint witness with an interior distribution or a loose approximate candidate
  • do not waive a failed sum, nonnegativity, marginal, or attainment constraint by changing the tolerance after the outcome
CC-005

FalsifierEvidence shows that an eligible E2 dataset was inspected, selected against, or analyzed before the contract was frozen, or that the stated executable conformance checks do not exist, or that a conforming dataset has been collected while the claim remains marked untested.

ConsequenceNARROW

Forbidden rescues

  • do not relabel pre-freeze exposure as “not inspection” after learning its outcomes
  • do not reissue a changed contract after data inspection and call it preregistered
  • do not cite the synthetic rehearsal as a conforming empirical dataset
CC-006

FalsifierReplaying the bound dry-run artifacts fails to reproduce the stated 66 conforming rows, zero validator violations, negative-control behavior, or the recorded n=22 null and cost calculation, or shows that a real E2 run was represented as the synthetic rehearsal.

ConsequenceREJECT

Forbidden rescues

  • do not substitute a repaired corpus, changed mechanism, or updated harness for the bound dry run
  • do not report a partial replay or a result under changed criteria as confirmation of the original rehearsal
REL-001

FalsifierAfter exactly 360 randomized eligible practitioners and a passing missingness gate, the upper endpoint of the frozen two-sided 95% Newcombe/Wilson interval for the Condition-B minus Condition-A HOLD risk difference is below +15 percentage points.

ConsequenceREJECT

Forbidden rescues

  • do not redefine success as engagement, aesthetics, confidence, perceived sophistication, comprehension, reading time, qualitative enthusiasm, or another secondary outcome
  • do not change the 15-point threshold, E2 study or pair, task threshold, primary outcome, interval method, sample target, stopping rule, or exclusions after outcomes are known
  • do not use a post-hoc subgroup, relaxed eligibility screen, missingness filter, top-up, or rerun to rescue an unfavorable primary result
  • do not generalize a simulated task result to actual deployment, broad practitioner relevance, customer value, adoption, or safety
MC-001

FalsifierA qualifying public evaluation meeting criteria v1 and published on or before 2026-08-27 is shown to be absent from this corrected census, or an examined row is shown to misreport its source such that recomputed N, M, or K differ from the stated 20/14/5.

ConsequenceREJECT

Forbidden rescues

  • do not reinterpret "joint statistic" after a counterexample appears in order to keep the stated counts
  • do not reclassify an examined row without recording the change and its reason in the census correction history
  • do not treat the bounded-search disclaimer as license to ignore a demonstrated miss — a miss rejects the current counts and the corrected census must state new ones
  • do not cite readership, links, or reuse of the census as evidence of its accuracy
MC-002

FalsifierRecomputing from the bound, hash-verified file yields any count different from the expected block, or the file at the pinned ref no longer matches the recorded sha256, or the released columns are shown not to be the labelled systems' verdicts.

ConsequenceREJECT

Forbidden rescues

  • do not substitute a different subset, column set, or harm-level filter to preserve the numbers
  • do not fold borderline prompts into either denominator after seeing the results
  • do not recast this subset arithmetic as a population estimate, with or without an interval, if the primary counts are challenged
  • do not cite the release-recomputed-to-plug-in ratio without the selection caveat that scopes it
MC-003

FalsifierA joint law with the stated marginals achieves an all-miss rate outside the interval; an exact finite catch-set arrangement lies outside the stated finite grid or exclusive-coverage bounds; or recomputing from the bound BELLS file yields any count different from the expected block.

ConsequenceREJECT

Forbidden rescues

  • do not change the composition rule after the fact to preserve the direction of the result
  • do not restate "identified only up to" as "estimated to be" — a bound is not an estimate
  • do not quote where a release-recomputed value sits inside the interval as a score or a percentage of a gap closed; the endpoints come from adversarial couplings with no detector-behavioural content
  • do not cite the zero-exclusive-coverage results as a vendor ranking, a causal attribution, or evidence about any product outside this stratum
  • do not fold the borderline stratum into either denominator to change any figure
MC-004

FalsifierRecomputing from the bound, hash-verified files yields any count different from the expected block, or any pinned file no longer matches its recorded sha256, or any recomputed quantity that overlaps the release's printed metrics disagrees with the printed value, or the committed blocked columns are shown not to be the named guards' verdicts.

ConsequenceREJECT

Forbidden rescues

  • do not fold text and image strata — or harmful and benign files — into pooled denominators to move any figure; the strata are bound exactly as released
  • do not substitute verdicts from the release's other run directories (adaptive_run, carrier sweeps, rendering probes) to preserve a number bound to full_run
  • do not reinterpret ShieldGemma 2's deterministic text passes as missing data to shrink a denominator after seeing the results
  • do not cite the benign-image 250/250 union without attributing it to Llama Guard 3 Vision's 250/250 column, and do not cite the harmful strata without the benign strata
  • do not call an OR of the harness-normalized native labels a shared-event catch union, stack safety result, or three-independent- guard finding unless a source-defined event translation is added
  • do not recast these file counts as population estimates, with or without an interval, if the primary counts are challenged
AF-001

FalsifierThe report states an artifact share or comparison group under which the relative reduction falls outside [30%, 43%]; or a published figure differs from the transcription; or the closed-form bound disagrees with a direct sweep over feasible shares.

ConsequenceNARROW

Forbidden rescues

  • do not widen the stated interval after the fact to accommodate a figure that falls outside it
  • do not restate the bound as an estimate of the artifact-conversation rate
  • do not present the Description-rises / Discernment-falls pattern as causal, or as evidence that artifacts reduce evaluation rather than relocate it
  • do not repackage the framework's behaviour taxonomy into any commercial or certification artifact; the licence is NonCommercial ShareAlike
GA-001

FalsifierA source-level review of the bound thesis shows that it does not state soundness as a ternary relation over the whole parse → canonicalize → digest pipeline, or does not identify Ghost-Ark as the stated S2 Lab research artifact.

ConsequenceNARROW

Forbidden rescues

  • do not replace a whole-pipeline claim with a canonicalizer-only claim after the condition fires
  • do not treat verifier acceptance or a pathology-free sample as evidence of semantic safety, authorization, or correctness
GV-001

FalsifierA local run from the bound revision lacks the stated seven scenes, deterministic sample-row computations, or Evidence Mode for inspection.

ConsequenceNARROW

Forbidden rescues

  • do not count hard-coded or mocked screenshots as deterministic computations
  • do not recast a local-only source project as a hosted interactive demo
GCE-001

FalsifierThe bound superseding taxonomy no longer classifies the coefficient family as deprecated legacy, or no longer supports the stated dependence-aware supersession.

ConsequenceNARROW

Forbidden rescues

  • do not treat the existence of legacy code or UI as evidence that its framing remains current theory
  • do not substitute a generic preference for newer methods for a recorded theoretical supersession
SITE-001

FalsifierRecomputing the listed palette pairs under WCAG contrast rules finds any claimed AA pass absent, or the recorded QA does not include the stated rendering, mode, and interaction checks.

ConsequenceNARROW

Forbidden rescues

  • do not replace a failed palette pair with another color pairing or viewport and retain the original claim
  • do not treat visual preference or casual inspection as a substitute for WCAG computation or the listed QA checks
SITE-002

FalsifierA controlled violation of a required registry field, bound-evidence trigger, or review window reaches a successful named CI workflow, or the workflow no longer runs on pushes to main and weekly.

ConsequenceREJECT

Forbidden rescues

  • do not cite a normal green run as evidence that violations are rejected
  • do not count a manual review or a generated-page drift check as automatic schema, trigger, or freshness enforcement
  • do not treat a warning-only or report-only job as a build failure
E3-001

FalsifierRecomputing from the committed observation rows yields any quantity different from the expected block beyond 1e-12, or the observed joint miss is shown to lie outside the Fréchet interval its own marginals allow, or a quoted prediction in the bound run report is shown to have been edited after the outcome was visible.

ConsequenceREJECT

Forbidden rescues

  • do not re-threshold, re-draw, or re-pool after the outcome and report the new numbers as this run
  • do not re-run the bootstrap under a different seed or B and report the resulting interval as this one
  • do not restate the failed primary prediction, narrow it, or drop it from the record
  • do not treat E3B as a re-run, a correction, or a replacement of this result
E3B-001

FalsifierRecomputing from the committed observation rows yields any quantity different from the expected block beyond 1e-12, or a quoted prediction in the bound run report is shown to have been edited after the outcome was visible, or E3B is represented anywhere in this repository as a re-run, correction, or replacement of E3.

ConsequenceREJECT

Forbidden rescues

  • do not re-threshold, re-draw, or re-pool after the outcome and report the new numbers as this run
  • do not present E3B as a repair of E3, or E3 as superseded by it
  • do not restate either failed prediction, narrow it, or drop it from the record
  • do not treat the zero-width interval as a measured absence of dependence
E6-001

FalsifierRecomputing from the committed rows yields any registered quantity different beyond 5e-5, or the pinned kernel bytes are shown not to reproduce all twelve printed Table 3 numbers to the printed decimal, or GDP is shown not to be monotone in one of the five parameters over its published interquartile range, or the measured joint median of 8.6 is shown to lie outside the coupling range this claim reports, or the extraction is shown to have altered the kernel's arithmetic rather than only re-instantiating its own factory source.

ConsequenceREJECT

Forbidden rescues

  • do not re-pin to a different kernel build after seeing an outcome and report the new numbers as this run
  • do not change the discretisation, its weights, or the hold-out values after the fact and present the resulting range as this one
  • do not describe E6 as preregistered, or any figure in it as a prediction that held
  • do not restate the unresolved site-copy estimand as though the public record settled which summary it names
  • do not carry any number here across to a guardrail, a classifier, or E2
E7B-001

FalsifierRecomputing from the committed rows yields any registered quantity different beyond 1e-9, or the preregistration's commit is shown not to precede the results commit, or any of the nine hold-out score files is shown to have been retrieved or inspected before f646e136, or the operating-point rule applied in experiments/e7b/run/measure.py is shown to differ from the one PREREG.md states, or an observed joint miss is shown to lie outside the Fréchet interval its own marginals allow.

ConsequenceREJECT

Forbidden rescues

  • do not re-threshold, re-calibrate, change the 5% budget, or change the direction of the operating-point rule and report the new numbers as this run
  • do not lower the six-pair power floor, or raise it, after the fact
  • do not pool the sixteen read judges of E7 and its exploratory record into this panel and report the combined figure as this result
  • do not present this as a repair, correction or replacement of the void E7, and do not present E7's frontier judges as still available for a preregistered claim
  • do not restate a prediction, narrow it, or drop it from the record
  • do not carry these numbers to E2's guards, to a deployed stack, or to a vendor

Replay manifest

The exact commands that re-verify this record from a clean checkout. CI runs them on every push and weekly; nothing here requires trusting this page.

python scripts/verify_claims.py            # shape, bindings, triggers, freshness, coverage
python scripts/verify_census.py            # census rows, N/M/K recomputed, MC-001 coherence
python scripts/generate_ledger.py --check  # the ledger is generated, not hand-edited
python scripts/generate_modules.py --check # module pages match their registry
python scripts/generate_observatory.py --check  # this page matches the registry
python scripts/generate_missing_column.py --check  # campaign pages match the census
python scripts/generate_sitemap.py --check # sitemap lastmod matches git history
python scripts/verify_figures.py           # figure geometry, asserted to 1e-9
python scripts/mjgd_reference.py --test    # disclosure arithmetic identities
python scripts/reanalyze_bells_subset.py   # MC-002 recomputed from the hash-bound release
python scripts/reproduce_cc001.py          # clean-clone kernel reproduction + witnesses

Human review

Last owner review: 2026-09-14. These review events cannot be executed by CI, and the record says so instead of borrowing the executable triggers' credibility:

Related instrument, its own contract intact: CC-Framework's evidence cards ↗ — their manifest forbids aggregate scores and forbids rendering "not-run" as passing, pending, or healthy, so this observatory links them rather than re-plotting them.