Observatory — the registry as one field · generated from claims.yamlSchema v0.4 · 21 claims · 2 contradicted · 18 supported within scope · 1 publicly untested
Claim observatory
Every claim in this record, as a capsule: its proposition, its evidence
binding, its decay clock, the boundary it will not cross, and the condition that would
defeat it. A solid cyan rail marks a
public evidence binding; capsules without one carry a dashed neutral rail and an
attested chip that says so; the clock turns amber when a review window is
three-quarters spent and red when it lapses. There is no aggregate score here and
never will be — a quiet trigger is quiet, not "healthy", and these envelopes do not
average into one number.
CC-001
Supported within scope
cc.kernel.strict.frechet_bounds computes Fréchet–Hoeffding endpoint bounds for composed binary guardrail failure from declared marginals; for marginals (0.10, 0.10) it returns [0.0, 0.10] for the AND event and [0.10, 0.20] for OR.
CC-Framework's E1 dependence-evidence study established constructively that distributions with identical singleton failure rates (0.5) and identical pairwise overlaps (0.25) can differ in three-way failure probability — 0 under even parity, 0.25 under odd parity — so measuring every pair does not identify a three-guardrail stack. E1 is synthetic; its recorded decision is Narrow.
For marginals (0.10, 0.10) and the both-fail query, cc.kernel.strict identify() returns the sharp interval [0.0, 0.10] together with endpoint witness distributions that sum to one, are nonnegative, satisfy both marginal constraints, and attain each endpoint — verified to 1e-9 from a clean clone.
CC-Framework's E2 — the shared-item empirical guardrail pilot — is governed by a measurement contract frozen before any dataset was inspected; no conforming dataset has been collected, so E2 is untested, and the contract's conformance checks are executable.
CC-Framework's E2 measurement pipeline was rehearsed end to end on three synthetic guardrail mechanisms: 66 observation rows conform to the frozen E2 schema with zero validator violations, the pre-registered negative controls behave (the common-cause control attains pA(1-pA) to 1e-9), and the finding records the result as a null at n=22 with a quantified cost for real E2 (~1097 shared items for epsilon=0.05). E2 itself is not run.
HYPOTHESIS — Among English-reading professionals who made or materially prepared a go/hold recommendation for an AI-based system in paid work during the prior 12 months, an E2-derived dependence-aware uncertainty presentation will increase HOLD decisions by at least 15 percentage points relative to a stipulated-independence presentation of the same evidence.
attested — no public artifact — stated on the owner's responsibility
As of 2026-08-27, among 20 public guardrail evaluations meeting the Missing Column Census's frozen inclusion criteria (v1), 14 establish a shared item set and a common event definition; 5 provide one of the census's declared joint-evidence artifacts: a printed composition result or per-item outcomes that make one directly computable.
On the 170-prompt per-item subset publicly released with BELLS's 2025 misuse-detection evaluation (pinned commit below), the five specialized supervisors — Lakera Guard, Prompt Guard, LangKit, NeMo Guardrails, LLM Guard — have an OR-union of released binary verdict columns that flags 73 of the 82 prompts labelled harmful, leaving 9 all-miss (11.0%). The product of their individual miss rates is 3.5%; the release-recomputed all-miss rate is about 3.1× that independence plug-in on this file. The same OR-union flags 19 of the 50 benign prompts.
For k guardrails scored on a common item set at a common operating point under block-on-any composition, the rate at which every guard misses the same item is fixed by the published per-guard miss rates only up to the interval [max(0, Σp − (k−1)), min p], and every point of that interval is attained by some joint law with those marginals. Marginals prove non-degradation of a static OR composition relative to its best member, but, when that identified set is non-degenerate, they do not identify strictly positive incremental benefit. They also identify a positive lower bound on static benign union-flag rate when any member has a positive benign flag rate. On the BELLS 2025 partial per-item release — 82 prompts labelled harmful — the exact finite identified set is {0/82 … 12/82}; MC-002 recomputes 9/82 all-miss from the hash-verified released verdict file. Marginal catch counts plus the union identify only bounds on each guard's exclusive full-stack coverage; the registered leave-one-out unions identify the realized values: NeMo 18, Lakera 3, and LangKit, Prompt Guard, and LLM Guard 0 on this stratum.
On the per-item guard-verdict files publicly released with Multimodal Safeguard Bench's full_run (pinned commit below; six verdict files and the release's two printed-metric files hash-verified), the three guards — Llama Guard 4, Llama Guard 3 Vision, ShieldGemma 2 — have harness-normalized native `unsafe` labels whose static Boolean OR is 1 on 192 of 200 harmful-labelled text items (8 all-zero rows, 4.0%) and all 200 harmful-labelled image items (0 all-zero rows), while the same OR is 1 on 32 of 250 benign-labelled text items (12.8%) and every one of 250 benign-labelled image items. Llama Guard 3 Vision alone has a 1-bit on all 250 released benign image items. The complete 2^3 native-label pattern table, leave-one-out bit ORs, and unique bit contributions for each released stratum are bound below. The release prints per-guard metrics and two two-guard compositions; it prints no three-guard bit OR, no all-zero-row count, no leave-one-out table, and no pattern decomposition — and every recomputed quantity that overlaps what it does print is asserted equal to the printed value.
Anthropic's AI Fluency Index reports overall prevalences for the 11 of its framework's 24 behaviours that are directly observable in conversations, and percentage-point differences for conversations that produced an artifact, but not the artifact share and not which comparison group the differences are against. Those published numbers therefore do not identify the subgroup rates — but they do bound them. Under every feasible artifact share and under either reading of the comparison, fact-checking is between 30% and 43% less prevalent in artifact conversations, a far larger effect than the stated 3.7 percentage points suggests on a base rate of 8.7%. The three behaviours that rise in artifact conversations are Description behaviours; the three that fall are Discernment behaviours.
Ghost-Ark is a verifier and measurement harness for the provenance limits of AI-governance receipts — a research artifact of the S2 Lab, Penn State — whose stated thesis is that receipt soundness is a ternary relation Sound(C, Σ, P): a receipt identifies an execution only up to the kernel of its whole parse → canonicalize → digest pipeline, so soundness does not persist by default as the pathology alphabet grows or the consumer set widens.
Ghost Visualizer is a seven-scene React/Vite visual essay ("Why AI Safety Scores Lie") that computes marginals, OR composition, Fréchet–Hoeffding bounds, endpoint witnesses, and a receipt object from deterministic sample rows, with an Evidence Mode for inspection.
GCE — the Guardrail Composability Explorer — is a coursework MVP demo (AI-285) for toggling guardrails and observing composed behavior. Its front-door composability-coefficient framing is superseded by CC-Framework's dependence-aware partial identification, whose metric taxonomy classifies the older coefficient family as legacy/deprecated. GCE is preserved as intellectual lineage, not current theory.
This site's palette token pairs were computed to pass WCAG AA contrast (most AAA), and manual QA passes covered light/dark themes, desktop and mobile layouts, keyboard focus, reduced motion, and no-JS rendering, as recorded in DESIGN.md.
This site's claim registry is enforced in CI: verify_claims.py validates schema and bindings, executes the executable review triggers against live evidence, and fails the build when any claim passes its freshness window — on every push and weekly.
E3 — the first pilot this repository ran itself — scored two ungated classifiers on 400 harmful and 800 benign items and produced 2,400 committed observation rows. Its primary pre-registered prediction FAILED: excess joint miss was +0.0018 with a 95% bootstrap CI of [-0.00096, +0.00706], which includes zero. Both guards missed almost everything on this pool (0.9825 and 0.9625), so the Fréchet interval the two marginals allow is [0.9450, 0.9625] — 1.75 percentage points wide — and the observed joint miss of 0.9475 lies inside it. The prediction that it would lie inside HELD; the difficulty-stratification prediction was NOT COMPUTED and remains open.
E3B put the same two classifiers on the attack family they were built for — 400 real prompt injections — and produced a further 2,400 committed observation rows. Its primary prediction FAILED and so did the prediction the redraw existed to test. One guard missed nothing (0.0000, catching 400 of 400) and the other missed 0.3975, so the Fréchet interval the marginals allow is [0.0000, 0.0000] — zero points wide — the observed joint miss is exactly 0.0000, and the bootstrap CI is the degenerate [0, 0]. E3B was pre-registered to produce an interval wider than 10 percentage points; it produced one narrower than E3's. The prediction that the observed joint miss would lie inside the interval HELD, trivially.
REJECTED AS STATED on 2026-09-10. The permutation routine moves unequal-mass atoms. Only 16 of 1,296 constructions preserve the specified marginal weights. The advertised 2.12–18.07 fixed-marginal range is unsupported and is rejected. No replacement range is asserted. Historical assertion retained for audit, not an accepted current finding: E6 extracted the client-side model kernel that anthropic.com/institute/econ-scenarios ships (Turbopack module 303412, sha256 0a75763f…), gated it against the twelve printed numbers of Table 3 of the Anthropic Institute's Working Paper 2026-02, and evaluated it only at parameter vectors built from the five marginal quantiles that paper prints in Table 2. The kernel reproduces all twelve Table 3 numbers to the printed decimal, and GDP is strictly increasing in each of the five parameters over their published interquartile ranges. Holding the four parameters the note to Table 4 names at that note's values, the same five marginals admit a median 2030 GDP anywhere in [2.12, 18.07] percent above the no-AI path across the 1,296 rank-permutation couplings of their published quartiles: 9.88 at the comonotone corner, 8.32 under independence, against the 8.6 the paper obtained by running each of 3,259 respondents' own five-vector. The marginals fix an interval fifteen points wide, not a point, and the measured joint lies strictly inside it.
REJECTED AS STATED on 2026-09-10. Calibration uses the all-nine shared pool instead of each judge’s own pool, changing one threshold from 0.95 to 1.00. Pair evaluation also uses the global intersection. The preregistration-conforming confirmation claim is rejected; the original HELD labels are historical outputs. Historical assertion retained for audit, not an accepted current finding: E7B is the first non-degenerate preregistered measurement this repository has produced, and all five of its predictions held on data unread when they were fixed. Against a preregistration frozen at f646e136 before any hold-out byte was retrieved, six of nine published guardrail-judge score sets reached a 5% benign false-flag budget on their shared 96-item AgentDojo pool (21 injection goals, 75 benign) and yielded 15 non-degenerate pairs. Median excess joint miss was +0.2018 and mean +0.1852 (P1, P2 HELD); the median pair's observed joint miss sat 0.900 of the way to the Fréchet upper bound its own marginals allow (P3 HELD); every observed joint lay inside its interval (P4 HELD); and every one of the 15 pairs had an observed joint miss above the independence product (P5 HELD). On the modal pair two judges each miss about half the injection goals, independence predicts a 0.249 both-miss rate, and the observed rate is 0.476 — exactly the upper bound, meaning one judge's misses are a subset of the other's. Seven of the fifteen pairs sit at that bound.
The outer boundary of the record: everything these claims refuse to
support, collected in one place. A claim without a stated non-claim is a claim that has
not found its edge yet.
CC-001
does not certify any stacked system as safe
does not recover or estimate the unknown dependence
does not select a point inside the returned interval
CC-002
the manifest's existence does not validate the claims it scopes
CC-003
the parity construction demonstrates possibility, not frequency, in real guardrail stacks — no claim of empirical prevalence
nothing here validates the framework's practical or deployed value; "we validated CC" is prohibited at every rung of the repository's evidence ladder
the product-baseline failure directions observed in E1 are properties of the tested generators, not universal constants
CC-004
endpoint witnesses are feasible mathematical worlds, not observed systems
attainment does not select a point inside the interval or estimate the true dependence
CC-005
freezing a contract establishes discipline, not results; E2 remains untested until conforming observations exist
conformance is not a clean bill of health, and the contract says so
"we validated CC" is prohibited at every rung of the evidence ladder
CC-006
this is not E2 and does not move E2 off untested; the mechanisms are toy filters in the repository, not deployed guardrails
the observed dependence is a null at this scale, not evidence of independence or of any real-guardrail dependence
validating the instrument establishes nothing about the safety, representativeness, or deployment behavior of any real system
REL-001
no practitioner decision, E2-derived participant packet, consent record, recruitment, or study outcome exists yet
this does not establish that dependence-aware evidence changes real deployment decisions, is generally useful, or justifies its cost
this is not a customer-demand, payment, adoption, safety, or certification claim
MC-001
does not claim any evaluated guardrail stack performs poorly; an unfilled column is a reporting fact, not a performance finding
does not claim the unmeasured joint statistics would reveal dependence; measuring instead of assuming is the point
does not audit the quality of any per-system evaluation beyond the fields each row records
does not claim exhaustive coverage beyond the corrected, examined artifact set or coverage after the stated date
MC-002
the subset is author-selected with an unstated rule — nothing here estimates any system's true rate, and no confidence interval is offered because the sampled population is undefined
not a ranking or endorsement of any vendor; the released verdicts are at unstated default configurations
says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release
the release-recomputed-to-plug-in ratio describes this subset's arithmetic, not a general law of guardrail dependence
not an observed deployed five-guard stack; interpreting this static OR aggregation as a stack requires the separate full-exposure, parallel, fixed-operating-point assumptions
MC-003
no new mathematics is claimed; the bounds are Fréchet's and the lower endpoint is Bonferroni's
not a claim that guardrails in general fail together — this stratum cannot separate shared blind spots from prompt-difficulty heterogeneity
not a vendor evaluation; the released verdicts are at unstated default configurations, and one supervisor fires exactly once in the 170 released rows
the identified interval is what the marginals leave open, not a prediction about any deployed stack
leave-one-out unions identify only exclusive full-stack coverage; they do not identify pairwise or higher-order overlap, Shapley values, or causal contribution
says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release
MC-004
not a shared-event catch statistic: the common `blocked` bit is a harness normalization of distinct native predicates, and no source-defined translation to a common event E has been identified
not a population estimate — the items derive from HarmBench and XSTest under the release's own construction, and no interval is offered because no sampled population is defined
not a ranking, endorsement, or indictment of any guard or vendor; the verdicts are at native, unmatched operating rules
not an observed deployed three-guard stack; interpreting the OR aggregation as a stack requires the separate full-exposure, parallel, fixed-operating-point assumptions
says nothing about the release's carrier-prompt, adversarial-UAP, or cross-VLM runs, which are separate artifacts with their own contracts
the counts are about the committed verdict bytes; the release's own changelog documents that ShieldGemma 2's image scores are sensitive to the text-rendering stack, so nothing here predicts what any guard would do under a different rendering environment
the zero three-guard image all-zero-bit count is a counting fact about these 200 released image items, not evidence of general image attack safety
AF-001
not a criticism of the Index; the observability boundary and the correlational limits are the report's own declared statements
not a claim about any individual's competence, and not a measure of anyone's diligence
not evidence that artifact production causes less checking — the published marginals cannot separate a behaviour change from a measurement change
no access to the underlying conversations; this is arithmetic on a published summary
GA-001
a verifying receipt does not establish that the governed action was safe, authorized, or semantically correct
kernel collisions in real canonicalizers are demonstrated as possible, not as prevalent — the repository's E12 sample found 0 of 64 real payloads carrying any pathology class
not hardened for deployment; not post-quantum secure
GV-001
not an AI-safety certificate; no deployment-readiness claim
its own public-readiness review scores it for private serious-contact demo use; the site therefore links source rather than exhibiting its media as a flagship
GCE-001
supersession is a statement about theoretical framing, not about the correctness of GCE's code or its value as coursework
a green run verifies registry consistency and quiet triggers, never the truth of any claim's content
an UNDETERMINED run (exit 2) blocks the build because a source was never reached; it is not a finding about any binding, and must not be read or recorded as a failed check
triggers watch file content; semantic drift outside watched files remains a manual review event, and the ledger says so
E3-001
the 2,400 rows prove the instrument runs end to end; they do not prove it measures what the programme says it measures, and on this pool the identified set was nearly a point, so it measured almost nothing
not evidence that either classifier is good or bad; two research models, one pool, one operating point each
the null is a null at this scale on this pool, not evidence of independence and not evidence about any real guardrail's dependence
says nothing about E2, its three frozen guards, its pools, or its operating points
no vendor, product, deployed stack, or population is described
E3B-001
a zero-width identified set means the marginals already fixed the joint miss; it is not a measurement that the two guards fail independently
G1 catching 400 of 400 is not evidence that it is a good guardrail, and contamination is unverified rather than excluded
G2 missing 159 of 400 is not evidence that it is a bad one
2,400 further rows prove the instrument runs; on this pool the question was degenerate, so they measure almost nothing
the sharpening these two pilots suggest — that marginal-only reporting is uninformative in the middle of the marginal range and fully informative at its extremes — is a hypothesis and an engineering design gate, not a result of either pilot; neither run tested it
no vendor, product, deployed stack, or population is described
E6-001
Current disposition is REJECT; the historical non-claims below do not restore the rejected assertion. See /corrections/#e6-marginals.
no error in the source paper is claimed; Table 4 does the joint-preserving computation and does it correctly, and the kernel agrees with the paper wherever both speak
not a claim that the site's "GDP is 10% higher" figure was computed by composing marginals; two different summaries of the same survey both round to it, the public record does not say which, and what is recorded is that the estimand is not identified from the artifact
the coupling range is an inner bound under a declared discretisation, not a sharp Fréchet bound; in five dimensions the comonotone corner is attainable but the lower envelope is not a copula, and that optimisation was not solved
not evidence about the US economy, about AI's economic effects, or about whether any scenario is likely; it is a statement about what a model's inputs determine
not a guardrail measurement and not transferable to one; it is the same identification structure measured on a different object
says nothing about the competence or intent of the paper's authors or reviewers
not a claim that the model composes its five inputs wrongly: the elicited quantities are conditional — the explorer's adoption dial asks what share of the tasks AI *can* do people will use it for — so their product within one respondent is the chain rule and is exact. The coupling this claim varies is the joint distribution across the 10,980 respondents, which no chain rule fixes
the 243 rows are evaluations of someone else's model at published parameter vectors, not per-item measurements of a guard on an item; they are counted by the evidence ledger as rows this repository produced, and the ledger's single observation-row total must not be read as growth in guard-item measurement. E3 and E3B's 4,800 rows are that kind; these are not
E7B-001
Current disposition is REJECT; the historical non-claims below do not restore the rejected assertion. See /corrections/#e7b-pools.
the Fréchet bound applied here is not this repository's: arXiv:2607.22868v1 states it for an any-flag gate as a proposition, and that artifact's own ensemble_robustness.py reports the vulnerability as correlated across judges. That author found the bound, released the per-item scores that make this checkable, and named the correlation; E7B measures it against a preregistration and claims priority for neither
21 harmful items is a small pool; every rate moves in steps of 0.048, no confidence procedure was preregistered, and no point estimate here implies an interval
15 pairs drawn from 6 judges are not 15 independent observations, and no multiplicity control was preregistered or is claimed
one benchmark, one pool, one operating point per judge; AgentDojo injection goals are not a deployed threat model and these are research judges, not a shipped stack
the three excluded judges were excluded by a pre-stated rule, and their absence is not evidence about them
not evidence that any specific product, vendor or deployed guardrail stack has correlated failures
five predictions holding is not a general law: it is one preregistered result on one pool, and the interpretation it supports is that independence understated the joint here, not that it always will
Falsifiers and forbidden rescues
A falsifier can otherwise be evaded by changing the proposition after
observing the result. Each record fixes its defeat condition and consequence in advance,
then lists the reinterpretations that cannot keep it standing. An explicit []
means no meaningful post-falsification rescue applies.
CC-001
FalsifierA clean-clone execution of the bound module on marginals (0.10, 0.10) returns AND or OR bounds other than the recorded endpoints, beyond the stated tolerance.
ConsequenceREJECT
Forbidden rescues
do not substitute a different input, revision, or numerical tolerance after the failure
do not treat a passing wrapper or test harness as evidence for the stated outputs if direct execution disagrees
CC-002
FalsifierThe bound manifest no longer maps its public claims to validation lanes, supporting files, and explicit non-claims as stated.
ConsequenceREJECT
Forbidden rescues
do not call partial or undocumented mappings complete manifest coverage
do not substitute a different, unbound checklist or README after the manifest fails this condition
CC-003
FalsifierThe frozen E1 construction fails to retain the stated equal singleton and pairwise values while yielding the stated different three-way probabilities (0 under even parity and 0.25 under odd parity), or the bound study does not record the stated Narrow decision.
ConsequenceREJECT
Forbidden rescues
do not switch to different marginals, overlaps, or generators after failure and call it the same construction
do not turn failure of the concrete construction into a generic possibility claim without a new, stated construction
CC-004
FalsifierBound identify() execution fails any stated interval, witness feasibility constraint, marginal constraint, or endpoint-attainment assertion at the declared tolerance.
ConsequenceREJECT
Forbidden rescues
do not replace a failed endpoint witness with an interior distribution or a loose approximate candidate
do not waive a failed sum, nonnegativity, marginal, or attainment constraint by changing the tolerance after the outcome
CC-005
FalsifierEvidence shows that an eligible E2 dataset was inspected, selected against, or analyzed before the contract was frozen, or that the stated executable conformance checks do not exist, or that a conforming dataset has been collected while the claim remains marked untested.
ConsequenceNARROW
Forbidden rescues
do not relabel pre-freeze exposure as “not inspection” after learning its outcomes
do not reissue a changed contract after data inspection and call it preregistered
do not cite the synthetic rehearsal as a conforming empirical dataset
CC-006
FalsifierReplaying the bound dry-run artifacts fails to reproduce the stated 66 conforming rows, zero validator violations, negative-control behavior, or the recorded n=22 null and cost calculation, or shows that a real E2 run was represented as the synthetic rehearsal.
ConsequenceREJECT
Forbidden rescues
do not substitute a repaired corpus, changed mechanism, or updated harness for the bound dry run
do not report a partial replay or a result under changed criteria as confirmation of the original rehearsal
REL-001
FalsifierAfter exactly 360 randomized eligible practitioners and a passing missingness gate, the upper endpoint of the frozen two-sided 95% Newcombe/Wilson interval for the Condition-B minus Condition-A HOLD risk difference is below +15 percentage points.
ConsequenceREJECT
Forbidden rescues
do not redefine success as engagement, aesthetics, confidence, perceived sophistication, comprehension, reading time, qualitative enthusiasm, or another secondary outcome
do not change the 15-point threshold, E2 study or pair, task threshold, primary outcome, interval method, sample target, stopping rule, or exclusions after outcomes are known
do not use a post-hoc subgroup, relaxed eligibility screen, missingness filter, top-up, or rerun to rescue an unfavorable primary result
do not generalize a simulated task result to actual deployment, broad practitioner relevance, customer value, adoption, or safety
MC-001
FalsifierA qualifying public evaluation meeting criteria v1 and published on or before 2026-08-27 is shown to be absent from this corrected census, or an examined row is shown to misreport its source such that recomputed N, M, or K differ from the stated 20/14/5.
ConsequenceREJECT
Forbidden rescues
do not reinterpret "joint statistic" after a counterexample appears in order to keep the stated counts
do not reclassify an examined row without recording the change and its reason in the census correction history
do not treat the bounded-search disclaimer as license to ignore a demonstrated miss — a miss rejects the current counts and the corrected census must state new ones
do not cite readership, links, or reuse of the census as evidence of its accuracy
MC-002
FalsifierRecomputing from the bound, hash-verified file yields any count different from the expected block, or the file at the pinned ref no longer matches the recorded sha256, or the released columns are shown not to be the labelled systems' verdicts.
ConsequenceREJECT
Forbidden rescues
do not substitute a different subset, column set, or harm-level filter to preserve the numbers
do not fold borderline prompts into either denominator after seeing the results
do not recast this subset arithmetic as a population estimate, with or without an interval, if the primary counts are challenged
do not cite the release-recomputed-to-plug-in ratio without the selection caveat that scopes it
MC-003
FalsifierA joint law with the stated marginals achieves an all-miss rate outside the interval; an exact finite catch-set arrangement lies outside the stated finite grid or exclusive-coverage bounds; or recomputing from the bound BELLS file yields any count different from the expected block.
ConsequenceREJECT
Forbidden rescues
do not change the composition rule after the fact to preserve the direction of the result
do not restate "identified only up to" as "estimated to be" — a bound is not an estimate
do not quote where a release-recomputed value sits inside the interval as a score or a percentage of a gap closed; the endpoints come from adversarial couplings with no detector-behavioural content
do not cite the zero-exclusive-coverage results as a vendor ranking, a causal attribution, or evidence about any product outside this stratum
do not fold the borderline stratum into either denominator to change any figure
MC-004
FalsifierRecomputing from the bound, hash-verified files yields any count different from the expected block, or any pinned file no longer matches its recorded sha256, or any recomputed quantity that overlaps the release's printed metrics disagrees with the printed value, or the committed blocked columns are shown not to be the named guards' verdicts.
ConsequenceREJECT
Forbidden rescues
do not fold text and image strata — or harmful and benign files — into pooled denominators to move any figure; the strata are bound exactly as released
do not substitute verdicts from the release's other run directories (adaptive_run, carrier sweeps, rendering probes) to preserve a number bound to full_run
do not reinterpret ShieldGemma 2's deterministic text passes as missing data to shrink a denominator after seeing the results
do not cite the benign-image 250/250 union without attributing it to Llama Guard 3 Vision's 250/250 column, and do not cite the harmful strata without the benign strata
do not call an OR of the harness-normalized native labels a shared-event catch union, stack safety result, or three-independent- guard finding unless a source-defined event translation is added
do not recast these file counts as population estimates, with or without an interval, if the primary counts are challenged
AF-001
FalsifierThe report states an artifact share or comparison group under which the relative reduction falls outside [30%, 43%]; or a published figure differs from the transcription; or the closed-form bound disagrees with a direct sweep over feasible shares.
ConsequenceNARROW
Forbidden rescues
do not widen the stated interval after the fact to accommodate a figure that falls outside it
do not restate the bound as an estimate of the artifact-conversation rate
do not present the Description-rises / Discernment-falls pattern as causal, or as evidence that artifacts reduce evaluation rather than relocate it
do not repackage the framework's behaviour taxonomy into any commercial or certification artifact; the licence is NonCommercial ShareAlike
GA-001
FalsifierA source-level review of the bound thesis shows that it does not state soundness as a ternary relation over the whole parse → canonicalize → digest pipeline, or does not identify Ghost-Ark as the stated S2 Lab research artifact.
ConsequenceNARROW
Forbidden rescues
do not replace a whole-pipeline claim with a canonicalizer-only claim after the condition fires
do not treat verifier acceptance or a pathology-free sample as evidence of semantic safety, authorization, or correctness
GV-001
FalsifierA local run from the bound revision lacks the stated seven scenes, deterministic sample-row computations, or Evidence Mode for inspection.
ConsequenceNARROW
Forbidden rescues
do not count hard-coded or mocked screenshots as deterministic computations
do not recast a local-only source project as a hosted interactive demo
GCE-001
FalsifierThe bound superseding taxonomy no longer classifies the coefficient family as deprecated legacy, or no longer supports the stated dependence-aware supersession.
ConsequenceNARROW
Forbidden rescues
do not treat the existence of legacy code or UI as evidence that its framing remains current theory
do not substitute a generic preference for newer methods for a recorded theoretical supersession
SITE-001
FalsifierRecomputing the listed palette pairs under WCAG contrast rules finds any claimed AA pass absent, or the recorded QA does not include the stated rendering, mode, and interaction checks.
ConsequenceNARROW
Forbidden rescues
do not replace a failed palette pair with another color pairing or viewport and retain the original claim
do not treat visual preference or casual inspection as a substitute for WCAG computation or the listed QA checks
SITE-002
FalsifierA controlled violation of a required registry field, bound-evidence trigger, or review window reaches a successful named CI workflow, or the workflow no longer runs on pushes to main and weekly.
ConsequenceREJECT
Forbidden rescues
do not cite a normal green run as evidence that violations are rejected
do not count a manual review or a generated-page drift check as automatic schema, trigger, or freshness enforcement
do not treat a warning-only or report-only job as a build failure
E3-001
FalsifierRecomputing from the committed observation rows yields any quantity different from the expected block beyond 1e-12, or the observed joint miss is shown to lie outside the Fréchet interval its own marginals allow, or a quoted prediction in the bound run report is shown to have been edited after the outcome was visible.
ConsequenceREJECT
Forbidden rescues
do not re-threshold, re-draw, or re-pool after the outcome and report the new numbers as this run
do not re-run the bootstrap under a different seed or B and report the resulting interval as this one
do not restate the failed primary prediction, narrow it, or drop it from the record
do not treat E3B as a re-run, a correction, or a replacement of this result
E3B-001
FalsifierRecomputing from the committed observation rows yields any quantity different from the expected block beyond 1e-12, or a quoted prediction in the bound run report is shown to have been edited after the outcome was visible, or E3B is represented anywhere in this repository as a re-run, correction, or replacement of E3.
ConsequenceREJECT
Forbidden rescues
do not re-threshold, re-draw, or re-pool after the outcome and report the new numbers as this run
do not present E3B as a repair of E3, or E3 as superseded by it
do not restate either failed prediction, narrow it, or drop it from the record
do not treat the zero-width interval as a measured absence of dependence
E6-001
FalsifierRecomputing from the committed rows yields any registered quantity different beyond 5e-5, or the pinned kernel bytes are shown not to reproduce all twelve printed Table 3 numbers to the printed decimal, or GDP is shown not to be monotone in one of the five parameters over its published interquartile range, or the measured joint median of 8.6 is shown to lie outside the coupling range this claim reports, or the extraction is shown to have altered the kernel's arithmetic rather than only re-instantiating its own factory source.
ConsequenceREJECT
Forbidden rescues
do not re-pin to a different kernel build after seeing an outcome and report the new numbers as this run
do not change the discretisation, its weights, or the hold-out values after the fact and present the resulting range as this one
do not describe E6 as preregistered, or any figure in it as a prediction that held
do not restate the unresolved site-copy estimand as though the public record settled which summary it names
do not carry any number here across to a guardrail, a classifier, or E2
E7B-001
FalsifierRecomputing from the committed rows yields any registered quantity different beyond 1e-9, or the preregistration's commit is shown not to precede the results commit, or any of the nine hold-out score files is shown to have been retrieved or inspected before f646e136, or the operating-point rule applied in experiments/e7b/run/measure.py is shown to differ from the one PREREG.md states, or an observed joint miss is shown to lie outside the Fréchet interval its own marginals allow.
ConsequenceREJECT
Forbidden rescues
do not re-threshold, re-calibrate, change the 5% budget, or change the direction of the operating-point rule and report the new numbers as this run
do not lower the six-pair power floor, or raise it, after the fact
do not pool the sixteen read judges of E7 and its exploratory record into this panel and report the combined figure as this result
do not present this as a repair, correction or replacement of the void E7, and do not present E7's frontier judges as still available for a preregistered claim
do not restate a prediction, narrow it, or drop it from the record
do not carry these numbers to E2's guards, to a deployed stack, or to a vendor
Replay manifest
The exact commands that re-verify this record from a clean checkout.
CI runs them on every push and weekly; nothing here requires trusting this page.
python scripts/verify_claims.py # shape, bindings, triggers, freshness, coverage
python scripts/verify_census.py # census rows, N/M/K recomputed, MC-001 coherence
python scripts/generate_ledger.py --check # the ledger is generated, not hand-edited
python scripts/generate_modules.py --check # module pages match their registry
python scripts/generate_observatory.py --check # this page matches the registry
python scripts/generate_missing_column.py --check # campaign pages match the census
python scripts/generate_sitemap.py --check # sitemap lastmod matches git history
python scripts/verify_figures.py # figure geometry, asserted to 1e-9
python scripts/mjgd_reference.py --test # disclosure arithmetic identities
python scripts/reanalyze_bells_subset.py # MC-002 recomputed from the hash-bound release
python scripts/reproduce_cc001.py # clean-clone kernel reproduction + witnesses
Human review
Last owner review: 2026-09-14. These review
events cannot be executed by CI, and the record says so instead of borrowing the
executable triggers' credibility:
CC-001 — kernel semantics change outside the bound module file
CC-003 — E2 produces empirical results — the synthetic scoping here must be restated
CC-004 — witness API change outside the bound module file
CC-005 — a conforming dataset is collected — E2's status must be restated
CC-006 — real E2 data is collected — this claim is superseded by an E2 result
REL-001 — a public protocol revision is bound, a real E2 result activates or prevents activation, or participant recruitment begins
MC-001 — a new qualifying evaluation is published or reported, an author corrects a row, or a benchmark publishes a previously absent joint statistic
MC-002 — the authors release the full per-item dataset, state the subset's selection rule, or correct the released verdicts
MC-003 — a reader identifies an error in the stated Fréchet, finite-grid, or aggregate exclusive-coverage bounds, or a result is used outside the declared static full-exposure assumptions
MC-004 — the authors publish a corrected or extended full_run, print their own three-guard joint statistics, document a matched operating-point calibration, or publish a source-defined translation from all native predicates to one shared catch event
AF-001 — the report publishes the artifact share, states which comparison group the differences are against, corrects a figure, or expands the analysis to a different sample
GV-001 — hosted deployment added — the site may then link a live demo
GCE-001 — GCE repository resumed or reframed
SITE-001 — site redesign
SITE-002 — workflow triggers or cadence change
E3-001 — a third pilot is run on a pool where both guards' miss rates are intermediate
E3B-001 — the contamination status of either guard's training corpus becomes checkable
E6-001 — Anthropic publishes the model's output at the Table 2 median vector, releases the 3,259 five-vectors, or ships a kernel whose digest differs from the pinned bytes
E7B-001 — a pool with intermediate marginals and more than 21 harmful items becomes available, or the upstream repository's scores change at a new commit
Related instrument, its own contract intact:
CC-Framework's
evidence cards ↗ — their manifest forbids aggregate scores and forbids rendering
"not-run" as passing, pending, or healthy, so this observatory links them rather than
re-plotting them.