The Missing Column — a source-bound census · generated from census.yaml19 examined · 13 shared item/common-event basis (0 at documented matched thresholds with full exposure) · 4 heterogeneous joint-evidence artifacts (1 via data release)criteria v1 wording locked 2026-08-27
The missing column
Teams may deploy guardrails in stacks, but a static evaluation and a
deployed route are different objects. This bounded record tracks whether public
evaluations preserve the joint-evidence artifacts needed to characterize a stated
static composition — and its own headline is generated from a source file that
anyone can mechanically make false.
An illustrative benchmark table: four guardrails
with individual catch rates, and a final column for a declared
full-exposure composition, which is not reported.
Guard A
Guard B
Guard C
Guard D
THE STACK
91%
88%
94%
86%
not reported
The missing column. Illustrative —
the shape of the reporting gap, not any specific evaluation's numbers. Four
individual catch rates on one attack set say almost nothing about the fifth
cell: a static joint result for all four on the stated item set. The census
below records which bounded inventory rows contain that evidence and which do
not.
Per-system rates do not determine the static all-miss rate of a
declared full-exposure composition. Two guards
that each miss 10% of attacks can jointly miss anywhere from 0% to 10% — the
individual columns are compatible with every world in that interval, and
multiplying the rates silently assumes the one world where the guards' failures
are independent. The front-page instrument lets you
move through those worlds with both marginals pinned;
the flagship essay
carries the full argument with witnessed endpoints, and
module 002 shows that
even pairwise numbers cannot rescue the inference. What resolves it is not more
per-system precision. It is one more column, measured on the same items.
What the second guard actually adds
One static composition question is: among the items the first
guard missed, what does the second catch? That is residual coverage, and no set
of per-guard columns contains it. Sequential route risk requires additional
observations beyond this static table.
What does the second guard catch among what
the first missed? Illustrative. Both worlds keep every per-guard rate
identical; only the overlap of misses moves. The static all-miss count differs by
a factor of 8. No table of per-guard columns
distinguishes these worlds; the missing column does. This figure asserts its own
geometry in CI and is not evidence about any real guardrail pair.
The census
Every row binds to its primary source and records the same
fields; the classification enum is fixed; the counts are recomputed from the file
by verify_census.py
on every push. The literal inclusion wording is locked in repository history before
row classification; that is a reproducibility lock, not an independent preregistration.
Inclusion — all must hold
public_accesspublicly accessible without payment at a stable URL
multi_systemevaluates two or more distinct, separately attributable guardrail, safeguard, or moderation systems — commercial APIs, open-source guard models, or models deployed as safeguards around an LLM
common_settinguses a common or purportedly common evaluation setting — the same dataset or prompt set claimed for all systems
domaintargets safety or harm detection, jailbreak or prompt-injection detection, content moderation, policy enforcement, or PII detection around LLM systems
method_detailprovides enough methodological detail to classify its joint-statistic status
Excluded by rule
base-LLM safety leaderboards: a model's own refusal behavior is not a separately deployed guardrail system
single-system evaluations: nothing to combine
marketing pages without measurements
comparisons only across different datasets, tasks, or populations with no purported common setting
Not a joint statistic
Prose recommending that systems be combined ("defense in depth", "use multiple providers") is a deployment recommendation, not a measured joint statistic. An average across systems is not a joint statistic. A multi-model ensemble inside one product counts only if the artifact reports it as a combination of separately attributed systems.
The bounded search protocol — executed 2026-08-27
This census claims coverage of the artifacts found by the documented search below — not of everything in existence. A qualifying artifact the search missed is a standing falsifier of any "among N" statement and is added, with the miss recorded in the revision history.
guard model comparison Llama Guard ShieldGemma benchmark
moderation endpoint comparison OpenAI Perspective
guardrail evaluation 2026
Snowball: one hop from the references of included artifacts ·
Budget: up to ~18 candidate artifacts examined in the frozen pass; candidates surfaced but not examined are listed under unexamined_candidates and excluded from every count
The census result — regenerated from the source file
As of 2026-08-27, among 19 public guardrail evaluations meeting the frozen inclusion criteria (v1), 13 establish a shared item set and a common event definition. That is a shared-basis rung, not evidence of matched operating thresholds or full exposure. 4 provide one of the declared joint-evidence artifacts: a printed composition result or per-item outcomes from which one is directly computable.
The 4 is a heterogeneous discovery count, not one deployment estimand. It splits as: 2 printed — covers the evaluated stack · 1 printed — covers part of the evaluated set · 1 computable — per-item outcomes released.
The 13 is a ladder, not a verdict — 13 document a
shared item set and a common event definition · 12 have no stated
threshold mismatch · 0 document matched
operating thresholds together with full exposure.
A qualifying artifact this search missed, or a row shown to be
misclassified, changes these numbers — the criteria and the correction route are
below, and every change lands in the revision history.
The census: each public guardrail evaluation
examined, with its joint-statistic status in the final column.
A Holistic Approach to Undesired Content Detection in the Real World — Markov, Zhang, Agarwal, Eloundou, Lee, Adler, Jiang, Weng — OpenAI,
2022-08-05 ·
primary source ↗
Task
undesired-content detection and moderation taxonomy (hate, sexual, violence, self-harm)
Population
OpenAI's public evaluation set plus external test sets scored by both systems: Jigsaw (5,000), TweetEval (2,970 hate / 860 offensive), Stormfront (478), Reddit (5,000).
Per-system results
AUPRC per category and dataset — Table 3 (e.g. 0.9703 vs 0.8709 on sexual content)
Joint statistic
none found
Why this classification
The minimal qualifying case — exactly two systems on shared external sets — and still no error-overlap or joint statistic; its released evaluation set became the de-facto shared test bed later papers reuse.
Same items for all systems
yesboth systems scored on the same external datasets (Table 3)
Same event definition
yesAUPRC computed against common labels per dataset (Table 3)
Thresholds comparable
unstatedAUPRC is threshold-free; operating thresholds not compared
All systems saw all items
yesboth scored per dataset; no gating design
Per-item outcomes released
noevaluation dataset released (github.com/openai/moderation-api-release); per-item Perspective outputs not released
Union detection reported
nonone found in results tables or text
All-miss rate reported
nonone found
Pairwise intersections reported
nono overlap analysis anywhere in the paper
Residual coverage reported
nono such decomposition anywhere
Joint uncertainty reported
unstatednot recorded in this examination
Systems
OpenAI moderation model, Perspective API
Source passages
Table 3 — AUPRC, OpenAI model vs Perspective API across datasets
released data: github.com/openai/moderation-api-release
Corrections
none recorded yet — the row invites them
Checked
2026-08-27
llama-guard-2023
not published
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations — Inan et al. — Meta GenAI,
2023-12-07 ·
primary source ↗
Task
prompt and response safety classification (moderation)
Population
Meta's internal test set (3,497) plus OpenAI Moderation Evaluation (1,680) and ToxicChat (10,000).
Per-system results
AUPRC per test set — Table 2 (Llama Guard 0.945 internal-prompt vs OpenAI API 0.764, Perspective 0.728); adaptability in Table 4
Joint statistic
none found
Why this classification
The foundational guard-versus-API comparison inherited by nearly all later guard papers, with no overlap analysis — the reporting pattern the rest of this census keeps finding starts here.
Same items for all systems
yesshared public test sets; Azure excluded from AUPRC because it returns no scores
Same event definition
yesper-API outputs binarized to a common unsafe verdict ('1-vs-all', '1-vs-benign' adaptations stated)
Thresholds comparable
unstatedAUPRC used for scoreable systems; per-API adaptation described, no calibration
All systems saw all items
yeseach scored per test set; no gating
Per-item outcomes released
nomodel and code at github.com/facebookresearch/PurpleLlama; no per-item comparison outputs
AUPRC and F1 per dataset (Table 3); accuracy on SimpleSafetyTests (Table 4)
Joint statistic
Section 5 with Figures 2-3: an Exponential Weights learner routes each item to one of three named experts (LlamaGuardDefensive, LlamaGuardPermissive, NeMo43B-Defensive); the figures print the learner's cumulative-regret trajectory on the OpenAI Moderation stream, measured one prompt at a time.
Why this classification
A measured ensemble-of-guards result exists — but it is a routing ensemble over the authors' own three experts (3 of the 8 compared systems), printed as cumulative regret relative to the best expert in hindsight. Verified against the v2 full text: the figures plot the algorithm's curves only, not per-expert error traces. No union, all-miss, or overlap statistic between independent systems appears anywhere.
Same items for all systems
mixedshared test sets in Tables 3-4, but Perspective/GPT-4 SimpleSafetyTests numbers are quoted from the source paper rather than re-run
Same event definition
yessafe/unsafe verdicts against common labels per test set
Thresholds comparable
unstatedAUPRC plus F1 at unstated operating points
All systems saw all items
mixedauthors' systems and main baselines re-run; some baseline cells quoted from prior work
Per-item outcomes released
nodataset release announced as intent (~26k annotations); no per-system predictions published
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
nono error-overlap or confusion analysis between systems
Residual coverage reported
nono such decomposition
Joint uncertainty reported
mixedFigure 3 averages over 20 trials; no CIs on Table 3 metrics
Section 5 — 'Deploy safety LLM models AegisSafetyExperts as ensemble'
Figure 2 caption — 'Aegis learns to choose the best expert over the time horizon'
Figure 3 — 'EW with perturbation averaged over 20 trials'
Table 3 — per-system AUPRC/F1 on shared test sets
Corrections
none recorded yet — the row invites them
Checked
2026-08-27
wildguard-2024
not published
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs — Han, Rao, Ettinger, Jiang, Lin, Lambert, Choi, Dziri — Allen Institute for AI / UW / SNU,
2024-06-26 ·
primary source ↗
Task
prompt harmfulness, response harmfulness, and refusal detection; jailbreak filtering
Population
WildGuardTest (5,299 human-annotated items) plus ten public benchmarks (ToxicChat, OpenAI Moderation, AegisSafetyTest, SimpleSafetyTests, HarmBench prompt/response, SafeRLHF, BeaverTails, XSTest-Resp).
Per-system results
F1 per benchmark (Tables 3-4); XSTest-Resp F1 (Table 2); ASR/RTA with guards as jailbreak filters (Table 6)
Joint statistic
none found
Why this classification
Thirteen baselines on shared test sets, aggregate scores only; the released test data would let anyone compute the joint column by re-running the guards, but the paper neither computes nor releases it.
Same items for all systems
mixedsame test sets within each task, but the baseline subset differs by task capability (Tables 2-4)
Same event definition
yesper-task F1 against common labels
Thresholds comparable
unstatedF1 at native decision rules
All systems saw all items
mixednot all baselines evaluated on all tasks
Per-item outcomes released
noWildGuardTest data released (huggingface.co/datasets/allenai/wildguardmix); baseline per-item predictions not released
prompt-injection and jailbreak detection against hard negatives and benign text
Population
4,314 items (3,016 English, 1,298 non-English; 5.2% injections, 0.9% jailbreaks, 20.9% hard negatives, ~73% benign chats and documents); a blend of public and proprietary data, not fully released.
Per-system results
single PINT score (accuracy-style %) per system with test date — README scoreboard
Joint statistic
none found
Why this classification
Vendor-maintained scoreboard of single-system scores; proprietary items and vendor-run scoring make third-party joint statistics impossible without Lakera's cooperation.
Same items for all systems
mixedsame named benchmark, but scoreboard runs are dated months apart and the dataset is deliberately versioned (Goodhart-resistance)
Same event definition
yessingle PINT accuracy score against the benchmark's labels
Thresholds comparable
unstatedvendor defaults; single score per system
All systems saw all items
unstatedper-run dataset version not printed per row
Per-item outcomes released
noproprietary blend; per-system outputs not released
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
noscoreboard is single-system rows only
Residual coverage reported
nono such decomposition
Joint uncertainty reported
nopoint scores only on the scoreboard
Systems
Lakera Guard, AWS Bedrock Guardrails, Azure AI Prompt Shield, protectai/deberta-v3-base-prompt-injection-v2, Llama Prompt Guard 2 86M, Google Model Armor, Aporia Guardrails, Llama Prompt Guard
Table 1 — Optimal F1 / AU-PRC vs LlamaGuard, WildGuard, GPT-4, OpenAI Mod API
Abstract — '+10.8% higher average AU-PRC compared to LlamaGuard1'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27
upenn-multilingual-2024
not published
Benchmarking LLM Guardrails in Handling Multilingual Toxicity — Yang, Dan, Roth, Lee — University of Pennsylvania / Microsoft,
2024-10-29 ·
primary source ↗
Task
multilingual toxicity and safety detection plus jailbreak robustness
Table 2 — Moderation dataset F1 drops to 28.54-73.12 multilingual
conclusion — 'guardrails are still ineffective at handling multilingual toxicity'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27
guardbench-2024
not published
GuardBench: A Large-Scale Benchmark for Guardrail Models — Elias Bassani, Ignacio Sanchez — European Commission Joint Research Centre,
2024-11 (EMNLP 2024 main) ·
primary source ↗
Task
safe/unsafe classification of prompts and conversation utterances
Population
40 datasets (Table 1), including PromptsDE/FR/IT/ES (30,852 each), UnsafeQA (22,180), BeaverTails-330k test (11,088), Toxic Chat (5,083).
Per-system results
F1 (and recall for all-unsafe datasets) per model per dataset — Table 3
Joint statistic
none found
Why this classification
Thirteen guards, forty datasets, one harness — and every published number is a marginal. The released pipeline makes the joint column one re-run away, which the paper does not take.
Same items for all systems
yesone automated pipeline runs every model over identical datasets (Sec. 3.5)
Same event definition
yessafe/unsafe per dataset, common labels
Thresholds comparable
unstatedF1/recall at native decision rules
All systems saw all items
yespipeline design; no gating
Per-item outcomes released
nono prediction dumps; the library re-run 'saves the moderation outcomes' locally (Sec. 3.5)
Sec. 1 — 'comparing 13 models on 40 prompts and conversations safety datasets'
Table 3 — per-model F1/Recall
Sec. 3.5 — library 'saves the moderation outcomes in the specified output directory'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27
ibm-adversarial-prompt-2025
not published
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs — Zizzo et al. — IBM Research (NeurIPS 2024 Safe GenAI workshop),
2025-02-21 ·
primary source ↗
Task
jailbreak and adversarial prompt detection (input guardrails)
Population
11,387 in-distribution items (9,543 benign / 1,844 malicious) plus 10,703 out-of-distribution items, drawn from 27 datasets.
Fifteen defenses on identical prompt pools with released code — the census's clearest one-re-run-away case — and the paper that says no single guardrail suffices still publishes no statistic about more than one.
Same items for all systems
yesparaphrase, not quotation: the method evaluates each defense over the same in-distribution and out-of-distribution pools (single shared pool design; the paper does not print the words "same prompts")
Same event definition
yesmalicious/benign against common labels
Thresholds comparable
unstatedAUC plus point metrics at native rules
All systems saw all items
yessingle-defense evaluation over the full pools; no gating
Per-item outcomes released
nocode released (github.com/IBM/Adversarial-Prompt-Evaluation); per-item outcomes not published
“Off-loading the detection of all such vectors to one guardrail is a significant challenge” — discussion/conclusion. A recommendation to combine is recorded here because it is not a measured joint statistic.
conclusion — 'Currently, there is no one-size-fits-all solution'
Corrections
2026-08-27 — Adversarial verification flagged that the same-items evidence read like a quotation; reworded to say plainly it is a paraphrase of the shared-pool design.
Checked
2026-08-27
neuraltrust-2025
not published
Benchmarking Jailbreak Detection Solutions for LLMs — Ayoub El Qadi — NeuralTrust,
2025-04-30 ·
primary source ↗
Task
jailbreak detection (prompt classification)
Population
A private set of 400 prompts (200/200) and a public set of ~600 (JailbreakBench JBB-Behaviors 200 plus GuardrailsAI detect-jailbreak 100 jailbreak / 300 benign).
Per-system results
accuracy, F1, execution time per system per dataset (in-post tables)
Joint statistic
none found
Why this classification
Three commercial systems on shared sets, marginals only; the vendor's own product wins, by the largest margin on its own private set — provenance recorded, joint statistics absent either way.
Same items for all systems
yessame items per dataset across all three systems
Same event definition
yesjailbreak/benign against common labels
Thresholds comparable
unstatedvendor defaults
All systems saw all items
yesall three scored per set; no gating
Per-item outcomes released
nono data or outputs released; private set unpublished
results table — NeuralTrust 0.908 acc / 0.897 F1 (private) vs Bedrock 0.615 / 0.296
public-set table — Azure 0.610 acc / 0.510 F1
Corrections
none recorded yet — the row invites them
Checked
2026-08-27
llamafirewall-2025
present
LlamaFirewall: An open source guardrail system for building secure AI agents — Chennabasappa et al. — Meta,
2025-05-06 ·
primary source ↗
Task
agent security guardrails: prompt-injection detection and agent-alignment checking
Population
AgentDojo: 97 tasks × static traces from ten language models × two conditions (benign and injected) ≈ 1,940 traces; CodeShield is evaluated separately on CyberSecEval3-derived code completions.
Per-system results
ASR and utility per configuration — baseline .1763, PromptGuard .0753, AlignmentCheck .0289 (Sec. 4.3.2 table)
Joint statistic
Section 4.3.2 table: rows for baseline, each component alone, and "Combined (PromptGuard + AlignmentCheck)" — ASR .0175 and utility .4268 on the same AgentDojo traces.
Why this classification
A genuinely printed stacked row — both evaluated components combined on the same items, with the stack's residual attack success beside each component's. Scope caveats: both components are the same vendor's, from the framework under evaluation; CodeShield (the third framework component) is evaluated on a different dataset and is not in the combined row; no uncertainty is reported.
Same items for all systems
yesboth components scored on the same AgentDojo traces; each monitors its role's messages within the shared traces
Same event definition
yesattack success (ASR) and utility on the same task suite
Thresholds comparable
unstatedcomponents at their shipped operating points
All systems saw all items
yesboth run over the full trace set; the combined row composes them
Per-item outcomes released
noframework code released; per-trace outcomes not published
Union detection reported
nothe combined row reports residual attack success, not a union-of-catches statistic as such
All-miss rate reported
yesCombined (PromptGuard + AlignmentCheck) ASR .0175 is the fraction of attacks succeeding past both layers — an all-miss statistic on the attack set
Pairwise intersections reported
nono overlap decomposition beyond the combined row
Residual coverage reported
nocomponent-vs-combined printed; per-layer residual attribution not decomposed
Sec. 4.3.2 table — 'Combined (PromptGuard + AlignmentCheck)' ASR .0175, utility .4268
'combined setup using both PromptGuard 2 86M and AlignmentCheck powered with Llama 4 Maverick'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27
circleguardbench-2025
not published
CircleGuardBench — White Circle AI,
2025-05-07 ·
primary source ↗
Task
harmful-content blocking, jailbreak resistance, false positives, and runtime across 17 harm categories
Population
Public split gated behind a license at huggingface.co/datasets/whitecircle-ai/circleguardbench_public; item count not stated in the announcement or repo pages examined.
Per-system results
accuracy, recall, precision, F1, error ratio, avg runtime, and an 'integral score' — leaderboard tables
Joint statistic
none found
Why this classification
No joint statistic appears on the leaderboard or announcement; many comparability fields remain unstated (recorded as such), so this row counts in N but not in M.
Same items for all systems
unstateda single harness implies common items; not explicitly confirmed in examined pages
Same event definition
unstatedleaderboard macro-averages across metric types; per-metric event definitions not examined
Thresholds comparable
unstatednot addressed in examined pages
All systems saw all items
unstatednot addressed in examined pages
Per-item outcomes released
unstatednot found in examined pages; dataset gated
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
noleaderboard is single-system rows
Residual coverage reported
nono such decomposition
Joint uncertainty reported
unstatednot addressed in examined pages
Systems
White Circle guard models (2), Llama Guard, OpenAI Moderation, Google Moderation API, GPT-4o-mini judges (CoT/strict), ShieldGemma, PromptGuard
Source passages
blog — integral score combines 'accuracy and runtime performance'
repo — 17 harm categories; leaderboard with macro-average metrics
Corrections
none recorded yet — the row invites them
Checked
2026-08-27
unit42-genai-platform-2025
not published
How Good Are the LLM Guardrails on the Market? A Comparative Study on the Effectiveness of LLM Content Filtering Across Major GenAI Platforms — Huang, Bray, Rao, Ji, Hu — Palo Alto Networks Unit 42,
2025-06-02 ·
primary source ↗
Task
platform-level content-filter effectiveness: jailbreak blocking and benign false positives
Population
1,123 prompts — 1,000 benign from four public datasets plus 123 JailbreakBench-derived jailbreak prompts — run against each platform's filters over the same underlying model.
Per-system results
false-positive and false-negative counts and block percentages per platform — Tables 1-6
Joint statistic
none found
Why this classification
Per-platform input-filter catch rates near 53%, 91%, and 92% on the same 123 jailbreaks practically beg for a union row; the article never computes one. Platforms are anonymized but separately reported as three distinct systems. The census publishes a stricter named-products-only sensitivity instead of silently treating anonymization as immaterial.
Same items for all systems
yesran each platform's content filters on the same prompts
Same event definition
yesblock/pass against common benign/jailbreak ground truth
“output guardrails play a crucial complementary role by capturing additional harmful outputs” — conclusion (refers to output filters complementing model alignment, not combining vendors). A recommendation to combine is recorded here because it is not a measured joint statistic.
Source passages
Guardrail Providers on the Market — platforms anonymized as Platform 1, Platform 2, and Platform 3
Tables 1-2 — benign/jailbreak block counts per platform
Table 6 — model alignment vs output guardrail blocking
method — same prompts, all safety filters enabled per platform
Corrections
2026-08-27 — Fresh-context adversarial review identified anonymization as a material interpretive boundary. The frozen criterion's phrase "separately attributable" did not specify whether names were required, so the record does not rewrite that clarification into the freeze: it reports the primary treatment and a named-products-only sensitivity that removes this row and mechanically yields N/M/K = 18/12/4.
Checked
2026-08-27
bells-misuse-2025
present
The bitter lesson of misuse detection (BELLS evaluation and leaderboard) — Mariaccia, Segerie, Dorn — CeSIA,
2025-07-08 ·
primary source ↗ · archived copy ↗
Task
harmful-prompt and jailbreak detection: specialized supervisors versus generalist LLMs on a harm-severity by adversarial-sophistication grid
Population
Non-adversarial: 990 prompts (330 benign / 330 borderline / 330 harmful across 11 harm categories). Adversarial: ~4,165 prompts (narrative, syntactic, and PAIR-generated families; exact splits read from one pass of the paper body). Sources include JailbreakBench, HH-RLHF.
No joint statistic is printed. The released per-item verdict columns (170 non-adversarial + 8 adversarial prompts, 11 systems) make union, all-miss, and every intersection directly computable on that subset.
Why this classification
PRESENT solely through the partial per-item release: joint statistics are computable on the released ~3.5% subset (178 items, all-systems columns), which is a real but small exception — nothing joint is printed, and the headline population's outcomes remain unreleased. The subset sizes are stated wherever this row is cited.
Same items for all systems
yesestablished by table structure and by released per-item CSVs whose rows carry one prompt with all systems' verdicts as columns; not stated as a sentence
Same event definition
yescommon harm taxonomy and harm-level labels across systems
Thresholds comparable
unstatedvendor defaults, binary verdicts, no calibration
All systems saw all items
mixedunstated globally; verified on the released per-item subset where every row has all systems' verdicts
Per-item outcomes released
mixedpartial release: bells_leaderboard repo data/ non_adversarial_prompts.csv holds 170 prompts with 11 systems' binary verdicts as columns (plus 8 adversarial prompts) — about 3.5% of the population; the full dataset is available only by contacting the authors (Appendix 0.C).
Union detection reported
nonot reported
All-miss rate reported
nonot reported
Pairwise intersections reported
nonot reported anywhere; computable from the released subset
Residual coverage reported
nonot reported
Joint uncertainty reported
yesinterval bounds on BELLS scores and rates (Table 1); method under-specified
“combining multiple LLMs or using voting mechanisms” — leaderboard site FAQ (a voting scaffold is mentioned, unmeasured). A recommendation to combine is recorded here because it is not a measured joint statistic.
Source passages
Table 1 — 11 systems, BELLS score with intervals (marginals)
Appendix 0.C — 'raw data at our leaderboard GitHub repository'
BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems — Waibl (Univ. Graz; SPAR), Michalak (SPAR), Mariaccia (CeSIA),
2026-06-12 ·
primary source ↗ · archived copy ↗
Task
operational comparison (detection, FPR, latency, cost) of specialized guardrails and frontier LLMs used as supervisors: content moderation and jailbreak detection
Population
Content moderation input: 1,400 samples (100 per 11 harm categories plus 300 benign). Content moderation output: 1,300 synthetic pairs. Jailbreak: 6,406 samples from 720 base prompts across 13 jailbreak families, plus four external sets.
Per-system results
detection rate, FPR, latency, and cost per supervisor — Tables 8-10 (Appendix C.2); per-category Tables 13-14; Pareto frontier Figure 2
Joint statistic
none found
Why this classification
The largest and most recent supervisor comparison found — exactly 28 systems from 17 providers on identical workloads under one harness — and every published number is per-supervisor. Per-item outcomes exist privately by construction (the harness's own output format), so a release would flip this row to computable.
Same items for all systems
yes'on identical input/output and adversarial workloads' under a single evaluation harness
Same event definition
yesper-system result mappers binarize each vendor's output to the common verdict
Thresholds comparable
unstatedvendor defaults through the mappers; no calibration
All systems saw all items
yesone evaluation pass per supervisor over the workloads; no gating
Per-item outcomes released
nothe harness writes one JSON per prompt when a user runs it, but the authors' own per-item outcomes are not published — repo tree has no results directory (225 paths checked) and the leaderboard ships aggregate metrics JSONs only.
Union detection reported
nonone found; the Pareto frontier is over individual systems
All-miss rate reported
noclosest is per-dataset minima (detection dropping below 34.2%), which is a marginal
Pairwise intersections reported
noevery table row is a single supervisor
Residual coverage reported
nono such decomposition
Joint uncertainty reported
mixedone evaluation pass per supervisor (stated limitation) — no detection CIs; leaderboard JSON carries latency CIs only
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents — Liu, Xu, Wang, Jia, Gong — Duke University,
2025-10-01 ·
primary source ↗
Task
prompt-injection detection for web agents, text and image modalities
Population
Text: 991 malicious segments (601 with explicit instructions) plus 2,707 benign. Image: 2,022 malicious plus 948 benign.
Per-system results
TPR per attack category and FPR per benign category, per detector — Tables 5-7
Joint statistic
Section 5.3 ("if any detector flags the sample as malicious, the ensemble classifies it as malicious") with Ensemble-T and Ensemble-I rows in Tables 5-7, covering all base detectors of each modality; verified against the full text.
Why this classification
The cleanest printed union in the census: OR-ensembles over all base detectors of each modality, with both TPR and FPR reported on shared items. Members are mostly academic and open-weights detectors — but not entirely: Ensemble-I includes GPT-4o-Prompt, a closed commercial API prompted as a detector, and Ensemble-T's PromptArmor runs on GPT-4o. No purpose-built commercial guardrail product appears in either union.
Same items for all systems
yesall detectors of a modality run on the same malicious and benign sets (Tables 5-7)
Same event definition
yesmalicious/benign against common labels per modality
Thresholds comparable
unstateddetectors at native decision rules
All systems saw all items
yesfull per-modality sets for every detector; no gating
Per-item outcomes released
nodatasets and code released (github.com/Norrrrrrr-lyn/WAInjectBench); per-item result files not identified — outcomes reproducible by re-run
Union detection reported
yesEnsemble-T and Ensemble-I rows are OR-rule unions over all base detectors of the modality — TPR in Tables 5-6, FPR in Table 7
All-miss rate reported
nonot framed or reported as an all-miss rate (the union TPR's complement covers the member set, unstated)
Sec. 5.3 — 'if any detector flags the sample as malicious, the ensemble classifies it as malicious'
Tables 5-7 — TPR/FPR including Ensemble-T and Ensemble-I rows
Corrections
2026-08-27 — Adversarial verification (same day, pre-publication) caught the reason line claiming the ensemble members were "academic and open-source detectors" — false: Ensemble-I includes GPT-4o-Prompt, a commercial closed API, and PromptArmor uses GPT-4o. Corrected; this correction also withdrew a commercial-API clause from claim MC-001, per that claim's own forbidden rescues.
'Recall is the critical metric for safety applications'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27
truefoundry-2026
not comparable
Benchmarking LLM Guardrail Providers: A Data-Driven Comparison — Kashish Kumar — TrueFoundry,
printed 2026-05-19; a Wayback capture of the same URL exists from 2026-03-13, so the printed date is unreliable ·
primary source ↗ · archived copy ↗
Task
PII detection, content moderation, and prompt-injection detection through the TrueFoundry AI Gateway
Population
Three hand-curated, category-balanced sets of 400 samples each (~50/50 positive/negative), one per task; not released.
Per-system results
precision, recall, F1, accuracy, Wilson 95% CI, latency — one table per task
Joint statistic
none found
Why this classification
Identical items but per-provider ground truth: each provider is scored against its own expected_triggers labels, so the cross-provider ranking is not a same-event comparison and a union or all-miss statistic over these marginals would be ill-defined. The blog recommends combining providers for defense in depth without measuring any combination.
Same items for all systems
yes'evaluated against identical datasets through the TrueFoundry AI Gateway' — within each task; only the content-moderation table has multiple systems
Same event definition
no'Each sample carries per-provider ground truth labels (expected_triggers)' — providers are scored against different label sets
Thresholds comparable
unstatedprovider default configurations via the gateway
“or combine multiple providers for defense-in-depth” — Key Takeaways bullet. A recommendation to combine is recorded here because it is not a measured joint statistic.
Source passages
'Evaluation Methodology' — 'evaluated against identical datasets through the TrueFoundry AI Gateway'
Four enterprise providers, one ~80,000-prompt dataset, marginals only — and the Azure row is itself an undisclosed two-component stack (Content Safety plus Prompt Shield reported as one number, combination rule unstated): a stack reported as a marginal.
Same items for all systems
yesimplied rather than formally stated: one dataset, and the latency section reports each provider's processing time for 'the full dataset'
Same event definition
yesone benign/harmful ground truth from the 63/37 split; no per-provider labels
Thresholds comparable
no'default or near-default security configurations'; no calibration; false-positive rates span 6.9%-15.5%
All systems saw all items
yeseach provider processed the full dataset (latency section)
Per-item outcomes released
nono repository, data, or per-prompt outputs
Union detection reported
nonowhere
All-miss rate reported
nonowhere
Pairwise intersections reported
nonowhere
Residual coverage reported
nonowhere
Joint uncertainty reported
nono confidence intervals or error bars anywhere
Systems
AWS Bedrock Guardrails, Azure (Content Safety + Prompt Shield, reported as one row), Cisco AI Defense, Google Cloud Model Armor
Source passages
Table 1 — four provider rows (AWS, Azure, Cisco, Google), marginals only
'How did we compare the four providers?' — 'approximately 80,000 Dutch prompts'
'How fast are these guardrails in practice?' — 'processing times for the full dataset'
Executive summary — 'Cisco AI Defense achieved the best F1-score (0.845)'
Corrections
none recorded yet — the row invites them
Checked
2026-08-28
What this census does not claim
It does not claim any evaluated stack performs poorly — an empty column is a
reporting fact, not a performance finding.
It does not claim the unmeasured joint statistics would reveal dependence;
measuring instead of assuming is the entire point.
Its 4 is an inclusive discovery count of noninterchangeable artifacts — not an
all-miss rate, a stack-quality score, or a deployment conclusion.
It does not audit the quality of any per-system evaluation beyond the fields
each row records.
It covers the artifacts found by the documented bounded search — not
everything in existence. A qualifying artifact it missed falsifies the "among
N" statement and is added on discovery.
The boundary of the search
Examined and excluded
unsafebench-2024
UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images (Qu et al., CISPA — arXiv 2405.03486) — source ↗
Excluded: Fails frozen domain criterion 5: it evaluates image safety classifiers guarding image-generation platforms, not guardrails around LLM systems. Recorded loudly rather than quietly, because it is the strongest joint-statistic reporter the search found anywhere: Tables 5 and 12 print six OR-rule ensemble rows (four pairwise, one three-way, one five-way over the conventional classifiers, 'the image is unsafe if any classifier in the ensemble reports it'), as F1, with no all-miss or overlap decomposition. A census with a broader multimodal-moderation domain would count it PRESENT. This is also the artifact previously misremembered as "MSBench with two pairwise ensemble rows" — the acronym was wrong and the ensemble count is six, not two.
bells-framework-2024
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards (arXiv 2406.01364) — source ↗
Excluded: Framework and position paper: reviews safeguards narratively and instantiates one MACHIAVELLI-based detector baseline, but contains no separately attributable multi-system evaluation — fails criteria 2 and 3. The empirical BELLS artifacts are separate rows above.
JailDAM (VLM jailbreak detection with multiple baselines) — source ↗
AI Safety Directory: OpenAI Moderation API vs Azure AI Content Safety (2026) — provenance of its numbers unclear — source ↗
Lakera: Assessing GenAI Security Solutions in the Wild with PINT (event page) — source ↗
Maxim: Complete AI Guardrails Implementation Guide for 2026 (mentions Bifrost's five integrated guardrail providers) — source ↗
GuardBench leaderboard (extends the paper's comparison to newer guards) — source ↗
The missing row, specified
The fix is small enough to paste into a results table: union
detection and the all-miss rate over the same items, with the denominator and
event definition that make them meaningful. The
Minimum Joint Guardrail
Disclosure page carries the exact template, its preconditions, and a tested
reference implementation.
Corrections and revision history
Changes to this census are recorded here, in the page they
change — not only in a repository log. Row-level corrections
(3 recorded) live inside each row above. The canonical
correction policy states the
response and logging rules.
2026-08-27 — Census established. Criteria v1 inclusion wording was committed and is now locked in repository history; that lock is a reproducibility record, not an independent preregistration. The four starting cases entered as under_review pending primary-source examination.
2026-08-27 — First examination pass completed: 19 rows examined (one over the ~18 budget estimate, recorded here), 9 candidates excluded by rule, 15 surfaced candidates left unexamined. Pre-publication schema clarifications, made before any count was public: the ABSENT definition no longer presupposes comparability (that axis is counted by M); joint_scope value printed_pairwise_only generalized to printed_partial_stack. Corrections to the starting brief: the single BELLS row split into three artifacts (2024 framework — excluded as a framework paper; 2025 misuse-detection evaluation; 2026 BELLS-O), and the artifact remembered as "MSBench with two pairwise ensemble rows" is actually UnsafeBench (arXiv 2405.03486), which prints six OR-ensemble rows but evaluates image-platform safety classifiers and therefore fails frozen domain criterion 5 — recorded prominently under exclusions. LlamaFirewall added from the snowball hop.
2026-08-27 — Demonstration computed on the one per-item outcome release found by the bounded search: union, all-miss, and residual coverage for the five specialized supervisors in BELLS 2025's released 170-prompt subset, registered as claim MC-002 (file bound by commit and sha256, reproduction script in CI) and rendered on the disclosure page. Census counts are unchanged — MC-002 is this record's own computation, not something the examined artifact printed.
2026-08-27 — Fresh-context adversarial verification, run same day and before anything was published, independently reproduced N/M/K = 19/13/4 and every MC-002 count, then found: (1) the wainjectbench-2025 reason line falsely called its ensemble members "academic and open-source" — Ensemble-I includes GPT-4o-Prompt, a commercial closed API; the row is corrected, and claim MC-001's commercial-API clause is withdrawn rather than reinterpreted, exactly as its forbidden rescues require. (2) A generator bug was deleting the letter "n" from the rendered criteria and exclusion notes on the public page (a malformed regex character class) — fixed, with the second exclusion rule's YAML typing corrected and the verifier extended to type-check exclusion rules. (3) The frozen phrase "separately attributable" had not said whether a product name was required. Rather than calling a post-review clarification pre-frozen, the record now exposes both readings: the primary treatment counts consistently distinguished, anonymized systems as separately attributable; the named-products-only sensitivity excludes unit42 and mechanically yields N/M/K = 18/12/4. Census schema v2 carries that executed sensitivity and a frozen-wording lock; criteria v1 itself is unchanged. (4) The IBM row's same-items evidence is marked as paraphrase. The primary counts remain N=19, M=13, K=4.
2026-08-28 — M ontology correction, made before any merge to main. The single number M was doing work it had not earned: the frozen criteria define "comparable" as shared items plus a shared event definition only, but a reader meets that word expecting matched operating thresholds and full exposure. compute_counts now derives an M ladder mechanically from fields already recorded on every row — 13 shared basis, 12 with no stated threshold mismatch (ML6 states the mismatch), 0 documenting matched thresholds with full exposure — and the page renders all three rungs. The proposition template gained an inline gloss saying which reading M uses. No row's evidence or classification changed; this is a weakening of what the headline implies, not a rescue of it. The strongest rung being 0 is the honest headline result and is now printed as such.
2026-08-28 — Adversarial release audit found that the repair still used the word "comparable" in the primary proposition and described the criteria lock as though it proved pre-search timing. Both claims were too strong. The public proposition now names only the mechanical shared-item/common-event basis, calls the four a heterogeneous joint-evidence discovery count rather than one estimand, and says precisely what the repository-history lock proves. No row, classification, or count changed; this is a further narrowing of language before any merge or public post.
Correct this record
If a row misreads its source, a supposedly absent statistic
exists, or a qualifying evaluation is missing:
open an
issue ↗ or write to
bhavepranavwork@gmail.com.
A confirmed correction updates the row, the counts, and this history — being
corrected is the mechanism working, and correction credit is recorded in the row.
Benchmark authors: if you
retained one decision per item per system, the
minimum joint disclosure is one
table away — and this census reclassifies your row to
present the day you publish it.
Replay manifest
The exact commands that re-verify this census from a clean
checkout. Nothing on this page requires trusting this page.
python scripts/verify_census.py --counts # row shape + N/M/K recomputed from census.yaml
python scripts/generate_missing_column.py --check # this page matches the census file
python scripts/verify_figures.py # figure geometry, asserted to 1e-9
python scripts/mjgd_reference.py --test # the disclosure arithmetic, tested