Pranav Bhave — AI Assurance · Security Engineering · Evidence Systems
Four guardrails. Four scores. One missing column.
Guardrail evaluations score each detector and
usually omit what the declared stack misses together — the one number a deployment
decision actually rests on.
Illustrative: four guardrails with individual catch
rates, and an unreported final column for a declared static composition.
Guard A
Guard B
Guard C
Guard D
THE STACK
91%
88%
94%
86%
not reported
Illustrative numbers; the gap is the point —
no arithmetic on the four printed columns recovers the fifth. I measure what AI
guardrail stacks miss together, and build evidence systems that show exactly what
data can and cannot establish.
A benchmark can score each supervisor jointly
across detection, false positives, latency, and cost; a deployed stack still needs
union and all-miss rates on the same items.
20 evaluations examined5 preserve joint evidence0 at matched thresholds
MC-001: the 14 shared-item rows and the 14 no-joint-artifact rows are different sets; 0 document matched thresholds with full exposure — 2026-09-01 sensitivity.
Pranav Bhave · Penn State B.S. CS ’26 · Philadelphia
Start with your question
A research front door before the dossier.
The Missing Column — a source-bound census
Which public evaluations preserve the joint evidence,
and which leave it unmeasured?
A bounded, single-reviewer inventory of public guardrail
evaluations, examined against primary sources. Every row binds to its source with
quoted passages, a fixed classification, and its own correction history. The
inclusion wording is locked in repository history before any row was classified,
the headline is regenerated from a file anyone can mechanically make false, and a
qualifying artifact the search missed rejects the current counts rather than being
absorbed quietly.
Two guardrails each fail 10% of the time. How
often do they fail together? The common answer is 1%. The evidence permits
anywhere from 0% to 10% — multiplying the two rates quietly chose one
world out of many, and never said so. Choose for yourself:
Dragging changes only the overlap of the two failures. Each guard's marginal failure rate remains ten percent.
Fig. 02 — the same two scores, every world they permit.
Both guardrails fail 10% of the time in every position. Only the overlap moves. The cyan scale is
the full supported interval; independence is the amber tick — one named assumption-world, not the
answer. The endpoint worlds are witnessed, and the geometry of every frame is asserted in CI.
The full argument →
01
Current programs
Three programs, and no more. Evidence markers apply to the
technical project claims below; biographical facts are owner-attested unless linked.
Chips with a solid dot open public evidence.
Attested means stated on my
responsibility, dated, with no public artifact yet — provenance, not proof. Exact
bindings — commit SHAs, executable review triggers, falsifiers, forbidden rescues,
non-claims — live in the
evidence ledger, and each claim also has
its own page carrying the falsifier next to the
proposition. Everything older or superseded is in
the archive, with its succession recorded.
Research framework · Python · MIT · 2025 — present
CC-Framework
Treats composed-guardrail failure as a partial-identification problem: individual
rail failure rates are known, their joint dependence is not. Given marginal evidence
and declared dependence assumptions, it computes sharp Fréchet–Hoeffding bounds on
stacked-system failure with endpoint witnesses, and records the boundary of the
resulting claim — claim envelopes, evidence roles, receipts, decay semantics. Its
flagship computation and witnesses are re-reproduced from a clean clone by this
site's CI, weekly.
S2 Lab, Penn State · verification research · in development
Ghost-Ark
A verifier and measurement harness for the provenance limits of AI-governance
receipts. Its research claim: receipt soundness is a ternary relation —
Sound(C, Σ, P) — and a receipt identifies an execution only up to
the kernel of its whole parse → canonicalize → digest pipeline,
so soundness does not persist by default as pathologies and consumers grow. Collisions
are demonstrated as possible, not measured as prevalent — and the repo says so.
Non-claim — a verifying receipt does not
establish that the governed action was safe, authorized, or semantically correct.
Field artifact P₀ · evidence-bound interfaces · experimental
The record itself
This site is the third program: a claim registry rendered into a generated ledger,
bound to immutable commits, watched by executable review triggers, re-reproduced from
clean clones in CI. The claim observatory renders the registry as one field of
capsules, bindings, and decay clocks. The central hypothesis — that claim-bound
interfaces improve a reader's comprehension or calibration — is untested; the control
study is designed, not run.
Non-claim — aesthetic coherence is not
epistemic reliability; what this interface does to readers is an open question.
Ghost VisualizerA seven-scene visual essay, “Why AI Safety Scores Lie”: identical
marginal guardrail scores hiding shared misses, with computed bounds, endpoint
witnesses, and an inspectable receipt object. Runs locally; no hosted deployment
yet — and its own readiness review scopes it to private demos, so this stays a
source link.
AssayEarly-stage, private: enclave-attested media processing on AWS
Nitro Enclaves — which provenance claims can a verifier actually derive from
attestation, and which can it not. Attestation cannot establish a file's history
prior to the first attested operation; the boundary is the project.
Attested · 2026-08
— stated on my responsibility; no public artifact linked yet
Earlier workGCE and the superseded guardrail-composition line are preserved in
the archive as intellectual lineage, each with what it explored, what was wrong,
and what replaced it.
One experimental protocol — E₁ — not a universal grammar. It
fits controlled tests; theorems, historical claims, and calibration problems need other
test designs. The general form is the claim envelope, below.
Claim
What I say the system does. Named first, in writing.
Falsifier
The observation that would end the claim. If none exists, neither does the claim.
Control
The comparison that could embarrass it — run on purpose.
Non-claim
What this evidence will never support, stated before anyone asks.
Result
Two distributions sharing every singleton and pairwise moment; three-way failure of 0% in one and 25% in the other. Measuring every pair does not identify the triple.
Supported within scope — controlled synthetic · decision: Narrow
Lock the first four before the result slot fills, then let the result disagree. The result above is E₁'s: it survived, and it narrowed the claim rather than widening it — claim
CC-003 binds it.
A story can start a question. It cannot finish an answer. Untested is not
inconclusive: inconclusive means evidence arrived and failed to discriminate; untested
means the world has not answered yet. This page keeps the two apart — and keeps its own
rules scoped, as declared above.
Candidate C₁
The claim envelope
1 · Proposition
What exactly is asserted.
2 · Scope
Where it holds, and under which assumptions.
3 · Support
The evidence that bears on it, bound to immutable artifacts.
4 · Falsifier
The condition that defeats the proposition, with its fixed consequence.
5 · Test design
Control, comparator, calibration, proof, benchmark, or adversary.
What remains unestablished — and the post-falsification rescues the claim cannot use.
8 · Freshness
When, and on what trigger, it must be re-examined.
The evidence
ledger is generated from a registry that implements these fields — propositions,
separated provenance and status dimensions, immutable bindings, falsifier conditions
with fixed consequences, forbidden rescues, executable review triggers that watch the
bound evidence upstream, freshness windows, non-claims. Where a
trigger can't be executed by CI, the ledger says manual, plainly.
03
Research & training
Aug — Dec 2025 · Penn State
Independent research — LLM safety guardrail composition
Reinforcement, interference, and dependence-driven failure in composed LLM safety
systems, via probabilistic bounds and copula-family reasoning. Supervised by
Dr. Peng Liu (IST 496). CC-Framework is the primary deliverable.
Jan — May 2025 · Penn State
Research assistant — logic & verification tooling
Python CNF-conversion tooling for SAT-solving workflows, and evaluation of
LLM-assisted formalization of natural-language logic: translation reliability,
ambiguity handling, verification-readiness.
May 2026
Pennsylvania State University
B.S. Computer Science, minor in Cybersecurity. Coursework: computer security,
operating systems, algorithms, statistical inference, theory of computation.
Certified
AWS — Cloud Practitioner · AI Practitioner
Solutions Architect Associate (SAA-C03) and Security — Specialty in progress.
Certifications are owner-attested here: verification IDs are
available on request rather than published.
The current question · frozen before any dataset was inspected
Can dependence evidence be measured at feasible cost — before it has to be
assumed?
E2, the first empirical rung of the evidence ladder, is a shared-item measurement
contract frozen in advance; no conforming dataset has been collected, so it is untested —
and says so. The next artifact is a conforming observation set, or a documented failure
to obtain one. What is moving right now →
When marginals are not enoughThe flagship case: two 10% guardrails, the interval the evidence
actually supports, and the witnesses at both ends — with the kernel's recorded
output at a bound commit.
Noetic log 002 — the observatory ships, and gets gradedThe v0.4 shipping record rated without anesthesia, with a
one-month lookback contract whose unflattering conclusion is pre-registered.