Pranav Bhave

Cubits11 / Explore / 01

Same scores.
Different worlds.

Two guardrails each miss some cases. Their scores tell you how often. They do not usually tell you whether they miss the same cases.

Change what you know. See what remains possible.

What do the scores leave open?

Interactive construction · no measured system
Possible rate of both missingDerived · marginals alone
0%–10%

A range of compatible answers, not a confidence interval.

0%100%

The amber marker is 1%: what independence assumes.

One compatible joint distribution A constructed area-one square: neither misses 85%, A only 5%, B only 5%, both 5%.
Neither misses · plain field
85%
A only · cyan
5%
B only · gold
5%
Both miss · hatched
5%

Area represents probability. These four cells sum to one; they are a mathematical witness, not observed cases. Gold locates event B, not an evidential status.

Every point from 0% to 10% is compatible with two 10% marginal miss rates.

This static example works without JavaScript. Enable JavaScript to change the rates and assumptions.

01 / Set the individual scores
02 / Decide what you add
03 / Inspect a surviving world
Inspect the calculation

Let a and b be the marginal miss probabilities, and q the probability both miss.

max(0, a + b − 1) ≤ q ≤ min(a, b)

The witness is (1 − a − b + q, a − q, b − q, q), ordered as neither, A only, B only, both. Nonnegative cells and a sum of one give the interval. Independence adds q = a × b.

A supplied interval is intersected with these bounds. This is established two-event Fréchet arithmetic; the interface adds no new theorem. Inspect the calculation source.

What this does and does not establish

Exact input probabilities, a common population, and the same declared miss event. Displayed values are rounded to two decimal places. No sampling uncertainty is modeled. A static stack catches a case if either guard catches it, so its miss event is their intersection.

This is not a deployed route, a vendor comparison, or evidence that a real stack is safe or unsafe. User-entered joint intervals are hypothetical constraints, not measurements.

Inspect the bound’s registered scope · Inspect endpoint witnesses · Run a reproduction

A small cinema of evidence.

Existing films / choose one / sound off
02 / Constructed identified set

Thirteen worlds

Follow the possibilities left open by a partial description.

Scope and provenance ↗ · Download film

What this film does not claim (MC-003)
  • no new mathematics is claimed; the bounds are Fréchet's and the lower endpoint is Bonferroni's
  • not a claim that guardrails in general fail together — this stratum cannot separate shared blind spots from prompt-difficulty heterogeneity
  • not a vendor evaluation; the released verdicts are at unstated default configurations, and one supervisor fires exactly once in the 170 released rows
  • the identified interval is what the marginals leave open, not a prediction about any deployed stack
  • leave-one-out unions identify only exclusive full-stack coverage; they do not identify pairwise or higher-order overlap, Shapley values, or causal contribution
  • says nothing about adversarial prompts, which dominate the full evaluation and have no per-item release

Silent films with on-screen text. Each source manifest describes its thesis, construction, and non-claims. Read about the film method.

Follow the thread.

From intuition to inspection
Play / a different way in

Move the misses yourself.

Place individual misses on a shared sheet in the original Worldspace experiment.

Enter Worldspace ↗
Read / the evidence

Which column is missing?

Inspect the bounded census of public evaluations, then follow each row to its source.

Open the census ↗
Inspect / a changing record

What survived correction?

Read the corrections and historical work with their limits still attached.

Read the corrections ↗ · Archive
An answer becomes more interesting when you can see what would change it.