Pranav Bhave
The Missing Column — a source-bound census · generated from census.yaml 19 examined · 13 shared item/common-event basis (0 at documented matched thresholds with full exposure) · 4 heterogeneous joint-evidence artifacts (1 via data release) criteria v1 wording locked 2026-08-27

The missing column

Teams may deploy guardrails in stacks, but a static evaluation and a deployed route are different objects. This bounded record tracks whether public evaluations preserve the joint-evidence artifacts needed to characterize a stated static composition — and its own headline is generated from a source file that anyone can mechanically make false.

An illustrative benchmark table: four guardrails with individual catch rates, and a final column for a declared full-exposure composition, which is not reported.
Guard A Guard B Guard C Guard D THE STACK
91% 88% 94% 86% not reported
The missing column. Illustrative — the shape of the reporting gap, not any specific evaluation's numbers. Four individual catch rates on one attack set say almost nothing about the fifth cell: a static joint result for all four on the stated item set. The census below records which bounded inventory rows contain that evidence and which do not.
Inspect the census Publish the missing row Correct this record

Why the last cell cannot be inferred

Per-system rates do not determine the static all-miss rate of a declared full-exposure composition. Two guards that each miss 10% of attacks can jointly miss anywhere from 0% to 10% — the individual columns are compatible with every world in that interval, and multiplying the rates silently assumes the one world where the guards' failures are independent. The front-page instrument lets you move through those worlds with both marginals pinned; the flagship essay carries the full argument with witnessed endpoints, and module 002 shows that even pairwise numbers cannot rescue the inference. What resolves it is not more per-system precision. It is one more column, measured on the same items.

What the second guard actually adds

One static composition question is: among the items the first guard missed, what does the second catch? That is residual coverage, and no set of per-guard columns contains it. Sequential route risk requires additional observations beyond this static table.

Identical individual rates, different residual coverage Two panels, each showing an illustrative attack set of one thousand items. In both, guard A catches nine hundred and misses one hundred. The one hundred missed items are magnified into a second strip showing what guard B catches among them. In world one, B catches ninety of the hundred and ten remain uncaught. In world two, B catches only twenty of the same hundred and eighty remain uncaught. Guard B's overall rate is the same in both worlds; only the overlap of the two guards' misses differs, and per-guard rates do not report it. World i — the world independence assumes A catches 900 of 1,000 · misses 100 the 100 A missed, magnified ↓ B catches 90 of the 100 A missed · 10 remain uncaught in this static illustration World ii — a correlated-miss world — same marginals A catches 900 of 1,000 · misses 100 the 100 A missed, magnified ↓ B catches 20 of the 100 A missed · 80 remain uncaught in this static illustration
What does the second guard catch among what the first missed? Illustrative. Both worlds keep every per-guard rate identical; only the overlap of misses moves. The static all-miss count differs by a factor of 8. No table of per-guard columns distinguishes these worlds; the missing column does. This figure asserts its own geometry in CI and is not evidence about any real guardrail pair.

The census

Every row binds to its primary source and records the same fields; the classification enum is fixed; the counts are recomputed from the file by verify_census.py on every push. The literal inclusion wording is locked in repository history before row classification; that is a reproducibility lock, not an independent preregistration.

Inclusion — all must hold

  • public_accesspublicly accessible without payment at a stable URL
  • multi_systemevaluates two or more distinct, separately attributable guardrail, safeguard, or moderation systems — commercial APIs, open-source guard models, or models deployed as safeguards around an LLM
  • per_system_resultsreports separately attributable per-system quantitative results
  • common_settinguses a common or purportedly common evaluation setting — the same dataset or prompt set claimed for all systems
  • domaintargets safety or harm detection, jailbreak or prompt-injection detection, content moderation, policy enforcement, or PII detection around LLM systems
  • method_detailprovides enough methodological detail to classify its joint-statistic status

Excluded by rule

  • base-LLM safety leaderboards: a model's own refusal behavior is not a separately deployed guardrail system
  • single-system evaluations: nothing to combine
  • marketing pages without measurements
  • comparisons only across different datasets, tasks, or populations with no purported common setting

Not a joint statistic

Prose recommending that systems be combined ("defense in depth", "use multiple providers") is a deployment recommendation, not a measured joint statistic. An average across systems is not a joint statistic. A multi-model ensemble inside one product counts only if the artifact reports it as a combination of separately attributed systems.

The bounded search protocol — executed 2026-08-27

This census claims coverage of the artifacts found by the documented search below — not of everything in existence. A qualifying artifact the search missed is a standing falsifier of any "among N" statement and is added, with the miss recorded in the revision history.

Fixed query list

  • LLM guardrail benchmark comparison
  • guardrails benchmark providers
  • jailbreak detection benchmark comparison providers
  • content moderation API comparison benchmark LLM
  • prompt injection detection benchmark comparison
  • LLM safeguard evaluation multiple providers
  • guard model comparison Llama Guard ShieldGemma benchmark
  • moderation endpoint comparison OpenAI Perspective
  • guardrail evaluation 2026

Snowball: one hop from the references of included artifacts · Budget: up to ~18 candidate artifacts examined in the frozen pass; candidates surfaced but not examined are listed under unexamined_candidates and excluded from every count

The census result — regenerated from the source file

As of 2026-08-27, among 19 public guardrail evaluations meeting the frozen inclusion criteria (v1), 13 establish a shared item set and a common event definition. That is a shared-basis rung, not evidence of matched operating thresholds or full exposure. 4 provide one of the declared joint-evidence artifacts: a printed composition result or per-item outcomes from which one is directly computable.

The 4 is a heterogeneous discovery count, not one deployment estimand. It splits as: 2 printed — covers the evaluated stack · 1 printed — covers part of the evaluated set · 1 computable — per-item outcomes released.

The 13 is a ladder, not a verdict — 13 document a shared item set and a common event definition · 12 have no stated threshold mismatch · 0 document matched operating thresholds together with full exposure.

A qualifying artifact this search missed, or a row shown to be misclassified, changes these numbers — the criteria and the correction route are below, and every change lands in the revision history.

The census: each public guardrail evaluation examined, with its joint-statistic status in the final column.
Artifact Systems Items Same items THE STACK
A Holistic Approach to Undesired Content Detection in the Real World Markov, Zhang, Agarwal, Eloundou, Lee, Adler, Jiang, Weng — OpenAI · 2022-08-05 2 ~14,300 across external sets yes not published
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Inan et al. — Meta GenAI · 2023-12-07 4 15,177 across three test sets yes not published
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts Ghosh, Varshney, Galinkin, Parisien — NVIDIA · 2024-04-09 8 ~4,000 across four test sets mixed presentprinted — covers part of the evaluated set
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Han, Rao, Ettinger, Jiang, Lin, Lambert, Choi, Dziri — Allen Institute for AI / UW / SNU · 2024-06-26 14 5,299 new + ten public benchmarks mixed not published
PINT Benchmark (Prompt Injection Test) Lakera AI · 2024-04 (scores updated through 2025-08) 8 4314 mixed not published
ShieldGemma: Generative AI Content Moderation Based on Gemma ShieldGemma Team — Google · 2024-07-31 7 ~20,680 across four sets mixed not published
Benchmarking LLM Guardrails in Handling Multilingual Toxicity Yang, Dan, Roth, Lee — University of Pennsylvania / Microsoft · 2024-10-29 4 ~13,000 yes not published
GuardBench: A Large-Scale Benchmark for Guardrail Models Elias Bassani, Ignacio Sanchez — European Commission Joint Research Centre · 2024-11 (EMNLP 2024 main) 13 40 datasets; six figures of items in total yes not published
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs Zizzo et al. — IBM Research (NeurIPS 2024 Safe GenAI workshop) · 2025-02-21 15 22090 yes not published
Benchmarking Jailbreak Detection Solutions for LLMs Ayoub El Qadi — NeuralTrust · 2025-04-30 3 ~1,000 across two sets yes not published
LlamaFirewall: An open source guardrail system for building secure AI agents Chennabasappa et al. — Meta · 2025-05-06 2 ~1,940 AgentDojo traces yes presentprinted — covers the evaluated stack
CircleGuardBench White Circle AI · 2025-05-07 at least 7 named unstated in examined pages unstated not published
How Good Are the LLM Guardrails on the Market? A Comparative Study on the Effectiveness of LLM Content Filtering Across Major GenAI Platforms Huang, Bray, Rao, Ji, Hu — Palo Alto Networks Unit 42 · 2025-06-02 3 1123 yes not published
The bitter lesson of misuse detection (BELLS evaluation and leaderboard) Mariaccia, Segerie, Dorn — CeSIA · 2025-07-08 12 ~5,155 yes presentcomputable — per-item outcomes released
BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems Waibl (Univ. Graz; SPAR), Michalak (SPAR), Mariaccia (CeSIA) · 2026-06-12 28 9,106 core (+4 external sets) yes not published
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents Liu, Xu, Wang, Jia, Gong — Duke University · 2025-10-01 12 6,668 across both modalities yes presentprinted — covers the evaluated stack
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation Harsh, Sarmah, Pasquali — Domyn · 2026-04-10 14 79331 yes not published
Benchmarking LLM Guardrail Providers: A Data-Driven Comparison Kashish Kumar — TrueFoundry · printed 2026-05-19; a Wayback capture of the same URL exists from 2026-03-13, so the printed date is unreliable 5 1,200 (400 per task) yes not comparable
Inside AI Guardrails: a benchmark on enterprise LLM security Cristóbal Sendín — ML6 (contributors Vrancken, Wehkamp, Van Der Burght) · 2026-06-22 (updated 2026-06-24) 4 ~80,000 (approximate, as stated) yes not published

openai-holistic-2022

not published

A Holistic Approach to Undesired Content Detection in the Real World — Markov, Zhang, Agarwal, Eloundou, Lee, Adler, Jiang, Weng — OpenAI, 2022-08-05 · primary source ↗

Task
undesired-content detection and moderation taxonomy (hate, sexual, violence, self-harm)
Population
OpenAI's public evaluation set plus external test sets scored by both systems: Jigsaw (5,000), TweetEval (2,970 hate / 860 offensive), Stormfront (478), Reddit (5,000).
Per-system results
AUPRC per category and dataset — Table 3 (e.g. 0.9703 vs 0.8709 on sexual content)
Joint statistic
none found
Why this classification
The minimal qualifying case — exactly two systems on shared external sets — and still no error-overlap or joint statistic; its released evaluation set became the de-facto shared test bed later papers reuse.
Same items for all systems
yesboth systems scored on the same external datasets (Table 3)
Same event definition
yesAUPRC computed against common labels per dataset (Table 3)
Thresholds comparable
unstatedAUPRC is threshold-free; operating thresholds not compared
All systems saw all items
yesboth scored per dataset; no gating design
Per-item outcomes released
noevaluation dataset released (github.com/openai/moderation-api-release); per-item Perspective outputs not released
Union detection reported
nonone found in results tables or text
All-miss rate reported
nonone found
Pairwise intersections reported
nono overlap analysis anywhere in the paper
Residual coverage reported
nono such decomposition anywhere
Joint uncertainty reported
unstatednot recorded in this examination
Systems
OpenAI moderation model, Perspective API
Source passages
  • Table 3 — AUPRC, OpenAI model vs Perspective API across datasets
  • released data: github.com/openai/moderation-api-release
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

llama-guard-2023

not published

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations — Inan et al. — Meta GenAI, 2023-12-07 · primary source ↗

Task
prompt and response safety classification (moderation)
Population
Meta's internal test set (3,497) plus OpenAI Moderation Evaluation (1,680) and ToxicChat (10,000).
Per-system results
AUPRC per test set — Table 2 (Llama Guard 0.945 internal-prompt vs OpenAI API 0.764, Perspective 0.728); adaptability in Table 4
Joint statistic
none found
Why this classification
The foundational guard-versus-API comparison inherited by nearly all later guard papers, with no overlap analysis — the reporting pattern the rest of this census keeps finding starts here.
Same items for all systems
yesshared public test sets; Azure excluded from AUPRC because it returns no scores
Same event definition
yesper-API outputs binarized to a common unsafe verdict ('1-vs-all', '1-vs-benign' adaptations stated)
Thresholds comparable
unstatedAUPRC used for scoreable systems; per-API adaptation described, no calibration
All systems saw all items
yeseach scored per test set; no gating
Per-item outcomes released
nomodel and code at github.com/facebookresearch/PurpleLlama; no per-item comparison outputs
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
nono error-overlap analysis anywhere
Residual coverage reported
nono such decomposition anywhere
Joint uncertainty reported
unstatednot recorded in this examination
Systems
Llama Guard, OpenAI Moderation API, Perspective API, Azure AI Content Safety (binary comparison only)
Source passages
  • Table 2 — AUPRC: Llama Guard vs OpenAI Moderation API vs Perspective API on three test sets
  • baseline handling — 'Overall binary classification for APIs'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

aegis-2024

present

AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts — Ghosh, Varshney, Galinkin, Parisien — NVIDIA, 2024-04-09 · primary source ↗

Task
content safety moderation of prompts and responses
Population
AegisSafetyTest (1,199), OpenAI Moderation dataset (1,680), ToxicChat, SimpleSafetyTests (100).
Per-system results
AUPRC and F1 per dataset (Table 3); accuracy on SimpleSafetyTests (Table 4)
Joint statistic
Section 5 with Figures 2-3: an Exponential Weights learner routes each item to one of three named experts (LlamaGuardDefensive, LlamaGuardPermissive, NeMo43B-Defensive); the figures print the learner's cumulative-regret trajectory on the OpenAI Moderation stream, measured one prompt at a time.
Why this classification
A measured ensemble-of-guards result exists — but it is a routing ensemble over the authors' own three experts (3 of the 8 compared systems), printed as cumulative regret relative to the best expert in hindsight. Verified against the v2 full text: the figures plot the algorithm's curves only, not per-expert error traces. No union, all-miss, or overlap statistic between independent systems appears anywhere.
Same items for all systems
mixedshared test sets in Tables 3-4, but Perspective/GPT-4 SimpleSafetyTests numbers are quoted from the source paper rather than re-run
Same event definition
yessafe/unsafe verdicts against common labels per test set
Thresholds comparable
unstatedAUPRC plus F1 at unstated operating points
All systems saw all items
mixedauthors' systems and main baselines re-run; some baseline cells quoted from prior work
Per-item outcomes released
nodataset release announced as intent (~26k annotations); no per-system predictions published
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
nono error-overlap or confusion analysis between systems
Residual coverage reported
nono such decomposition
Joint uncertainty reported
mixedFigure 3 averages over 20 trials; no CIs on Table 3 metrics
Systems
LlamaGuardBase, NeMo43B, OpenAI Moderation API, Perspective API, GPT-4, LlamaGuardDefensive (authors'), LlamaGuardPermissive (authors'), NeMo43B-Defensive (authors')
Source passages
  • Section 5 — 'Deploy safety LLM models AegisSafetyExperts as ensemble'
  • Figure 2 caption — 'Aegis learns to choose the best expert over the time horizon'
  • Figure 3 — 'EW with perturbation averaged over 20 trials'
  • Table 3 — per-system AUPRC/F1 on shared test sets
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

wildguard-2024

not published

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs — Han, Rao, Ettinger, Jiang, Lin, Lambert, Choi, Dziri — Allen Institute for AI / UW / SNU, 2024-06-26 · primary source ↗

Task
prompt harmfulness, response harmfulness, and refusal detection; jailbreak filtering
Population
WildGuardTest (5,299 human-annotated items) plus ten public benchmarks (ToxicChat, OpenAI Moderation, AegisSafetyTest, SimpleSafetyTests, HarmBench prompt/response, SafeRLHF, BeaverTails, XSTest-Resp).
Per-system results
F1 per benchmark (Tables 3-4); XSTest-Resp F1 (Table 2); ASR/RTA with guards as jailbreak filters (Table 6)
Joint statistic
none found
Why this classification
Thirteen baselines on shared test sets, aggregate scores only; the released test data would let anyone compute the joint column by re-running the guards, but the paper neither computes nor releases it.
Same items for all systems
mixedsame test sets within each task, but the baseline subset differs by task capability (Tables 2-4)
Same event definition
yesper-task F1 against common labels
Thresholds comparable
unstatedF1 at native decision rules
All systems saw all items
mixednot all baselines evaluated on all tasks
Per-item outcomes released
noWildGuardTest data released (huggingface.co/datasets/allenai/wildguardmix); baseline per-item predictions not released
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
noeach guard evaluated in isolation
Residual coverage reported
nono such decomposition
Joint uncertainty reported
unstatednot recorded in this examination
Systems
WildGuard, Llama-Guard, Llama-Guard 2, Aegis-Guard-Defensive, Aegis-Guard-Permissive, MD-Judge v0.1, HarmBench-Llama, HarmBench-Mistral, BeaverDam-7B, LibrAI-LongFormer-harm, LibrAI-LongFormer-ref, OpenAI Moderation API, GPT-4 (gpt-4-0125-preview), keyword-based refusal detector
Source passages
  • Abstract — comparison against ten strong existing open-source moderation models
  • Table 3 — F1 across public prompt/response-harm benchmarks
  • Table 4 — F1 on WildGuardTest including the adversarial split
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

lakera-pint-2024

not published

PINT Benchmark (Prompt Injection Test) — Lakera AI, 2024-04 (scores updated through 2025-08) · primary source ↗

Task
prompt-injection and jailbreak detection against hard negatives and benign text
Population
4,314 items (3,016 English, 1,298 non-English; 5.2% injections, 0.9% jailbreaks, 20.9% hard negatives, ~73% benign chats and documents); a blend of public and proprietary data, not fully released.
Per-system results
single PINT score (accuracy-style %) per system with test date — README scoreboard
Joint statistic
none found
Why this classification
Vendor-maintained scoreboard of single-system scores; proprietary items and vendor-run scoring make third-party joint statistics impossible without Lakera's cooperation.
Same items for all systems
mixedsame named benchmark, but scoreboard runs are dated months apart and the dataset is deliberately versioned (Goodhart-resistance)
Same event definition
yessingle PINT accuracy score against the benchmark's labels
Thresholds comparable
unstatedvendor defaults; single score per system
All systems saw all items
unstatedper-run dataset version not printed per row
Per-item outcomes released
noproprietary blend; per-system outputs not released
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
noscoreboard is single-system rows only
Residual coverage reported
nono such decomposition
Joint uncertainty reported
nopoint scores only on the scoreboard
Systems
Lakera Guard, AWS Bedrock Guardrails, Azure AI Prompt Shield, protectai/deberta-v3-base-prompt-injection-v2, Llama Prompt Guard 2 86M, Google Model Armor, Aporia Guardrails, Llama Prompt Guard
Source passages
  • README scoreboard — Lakera Guard 95.22% (2025-05-02) through Llama Prompt Guard 61.82%
  • blog — dataset 'a blend of public and proprietary data'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

shieldgemma-2024

not published

ShieldGemma: Generative AI Content Moderation Based on Gemma — ShieldGemma Team — Google, 2024-07-31 · primary source ↗

Task
safety content moderation of user input and model output
Population
Internal ShieldGemma Prompt (4,500) and Response (4,500) sets plus OpenAI Moderation (1,680) and ToxicChat (10,000).
Per-system results
Optimal F1 and AU-PRC (Table 1); per-harm-type AU-PRC (Figure 3)
Joint statistic
none found
Why this classification
Marginals only; the limitations section recommends client-side threshold tuning, not layered guards, and no combination is measured.
Same items for all systems
mixedTable 1 coverage varies by baseline — not all baselines evaluated on all datasets
Same event definition
yesbinary safety verdicts against common labels per dataset
Thresholds comparable
unstatedOptimal F1 and AU-PRC; per-system operating points differ by construction
All systems saw all items
mixedas above — coverage varies by baseline
Per-item outcomes released
noaggregate metrics only; models on HuggingFace/Kaggle; test-data release unstated
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
nono overlap analysis
Residual coverage reported
nono such decomposition
Joint uncertainty reported
unstatednot recorded in this examination
Systems
ShieldGemma 2B, ShieldGemma 9B, ShieldGemma 27B, LlamaGuard, WildGuard, GPT-4, OpenAI Moderation API
Source passages
  • Table 1 — Optimal F1 / AU-PRC vs LlamaGuard, WildGuard, GPT-4, OpenAI Mod API
  • Abstract — '+10.8% higher average AU-PRC compared to LlamaGuard1'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

upenn-multilingual-2024

not published

Benchmarking LLM Guardrails in Handling Multilingual Toxicity — Yang, Dan, Roth, Lee — University of Pennsylvania / Microsoft, 2024-10-29 · primary source ↗

Task
multilingual toxicity and safety detection plus jailbreak robustness
Population
~13,000 items across seven datasets (ToxicChat 1,000; Aegis 1,199; Moderation 1,680; RTP-LX 999; PTP 5,000; MultiJail 315; XSafety 2,800) in ten-plus languages.
Per-system results
F1 per dataset and language (Table 2); FPR on Aegis
Joint statistic
none found
Why this classification
Four open guard models, fully shared items, marginals only — a clean instance of the pattern on a multilingual axis.
Same items for all systems
yessame items per dataset for every guardrail (Table 2 structure)
Same event definition
yescommon toxicity/safety labels per dataset
Thresholds comparable
unstatedF1 at native decision rules
All systems saw all items
yesall four guards scored per dataset; no gating
Per-item outcomes released
nono repository stated
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
nono overlap analysis
Residual coverage reported
nono such decomposition
Joint uncertainty reported
unstatednot recorded in this examination
Systems
LlamaGuard-2, LlamaGuard-3, Aegis-Defensive, MD-Judge
Source passages
  • Table 2 — Moderation dataset F1 drops to 28.54-73.12 multilingual
  • conclusion — 'guardrails are still ineffective at handling multilingual toxicity'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

guardbench-2024

not published

GuardBench: A Large-Scale Benchmark for Guardrail Models — Elias Bassani, Ignacio Sanchez — European Commission Joint Research Centre, 2024-11 (EMNLP 2024 main) · primary source ↗

Task
safe/unsafe classification of prompts and conversation utterances
Population
40 datasets (Table 1), including PromptsDE/FR/IT/ES (30,852 each), UnsafeQA (22,180), BeaverTails-330k test (11,088), Toxic Chat (5,083).
Per-system results
F1 (and recall for all-unsafe datasets) per model per dataset — Table 3
Joint statistic
none found
Why this classification
Thirteen guards, forty datasets, one harness — and every published number is a marginal. The released pipeline makes the joint column one re-run away, which the paper does not take.
Same items for all systems
yesone automated pipeline runs every model over identical datasets (Sec. 3.5)
Same event definition
yessafe/unsafe per dataset, common labels
Thresholds comparable
unstatedF1/recall at native decision rules
All systems saw all items
yespipeline design; no gating
Per-item outcomes released
nono prediction dumps; the library re-run 'saves the moderation outcomes' locally (Sec. 3.5)
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
nono overlap analysis
Residual coverage reported
nono such decomposition
Joint uncertainty reported
unstatednot recorded in this examination
Systems
Llama Guard, Llama Guard 2, Llama Guard Defensive, Llama Guard Permissive, MD-Judge, Toxic Chat T5, ToxiGen HateBERT, ToxiGen RoBERTa, Detoxify Original, Detoxify Unbiased, Detoxify Multilingual, Mistral-7B-Instruct v0.2, Mistral with refined policy
Source passages
  • Sec. 1 — 'comparing 13 models on 40 prompts and conversations safety datasets'
  • Table 3 — per-model F1/Recall
  • Sec. 3.5 — library 'saves the moderation outcomes in the specified output directory'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

ibm-adversarial-prompt-2025

not published

Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs — Zizzo et al. — IBM Research (NeurIPS 2024 Safe GenAI workshop), 2025-02-21 · primary source ↗

Task
jailbreak and adversarial prompt detection (input guardrails)
Population
11,387 in-distribution items (9,543 benign / 1,844 malicious) plus 10,703 out-of-distribution items, drawn from 27 datasets.
Per-system results
AUC, accuracy, F1, recall, precision — Table 2 (in-distribution), Table 3 (out-of-distribution)
Joint statistic
none found
Why this classification
Fifteen defenses on identical prompt pools with released code — the census's clearest one-re-run-away case — and the paper that says no single guardrail suffices still publishes no statistic about more than one.
Same items for all systems
yesparaphrase, not quotation: the method evaluates each defense over the same in-distribution and out-of-distribution pools (single shared pool design; the paper does not print the words "same prompts")
Same event definition
yesmalicious/benign against common labels
Thresholds comparable
unstatedAUC plus point metrics at native rules
All systems saw all items
yessingle-defense evaluation over the full pools; no gating
Per-item outcomes released
nocode released (github.com/IBM/Adversarial-Prompt-Evaluation); per-item outcomes not published
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
nosingle-defense evaluation only
Residual coverage reported
nono such decomposition
Joint uncertainty reported
unstatednot recorded in this examination
Systems
Random Forest, BERT, DeBERTa, GPT2, ProtectAI v2, Llama-Guard 2, LangKit Injection Detection, SmoothLLM, Perplexity Threshold, OpenAI Moderation, Azure AI Content Safety, NeMo-inspired Input Rail, LangKit Proactive Defence, Vicuna-13b refusal, Granite Guardian 3.0
Combination prose
“Off-loading the detection of all such vectors to one guardrail is a significant challenge” — discussion/conclusion. A recommendation to combine is recorded here because it is not a measured joint statistic.
Source passages
  • Table 2 — DeBERTa 0.996 AUC vs ProtectAI v2 0.569 AUC (in-distribution)
  • Table 3 — Perplexity 0.824 AUC / 0.001 F1 (out-of-distribution collapse)
  • conclusion — 'Currently, there is no one-size-fits-all solution'
Corrections
  • 2026-08-27 — Adversarial verification flagged that the same-items evidence read like a quotation; reworded to say plainly it is a paraphrase of the shared-pool design.
Checked
2026-08-27

neuraltrust-2025

not published

Benchmarking Jailbreak Detection Solutions for LLMs — Ayoub El Qadi — NeuralTrust, 2025-04-30 · primary source ↗

Task
jailbreak detection (prompt classification)
Population
A private set of 400 prompts (200/200) and a public set of ~600 (JailbreakBench JBB-Behaviors 200 plus GuardrailsAI detect-jailbreak 100 jailbreak / 300 benign).
Per-system results
accuracy, F1, execution time per system per dataset (in-post tables)
Joint statistic
none found
Why this classification
Three commercial systems on shared sets, marginals only; the vendor's own product wins, by the largest margin on its own private set — provenance recorded, joint statistics absent either way.
Same items for all systems
yessame items per dataset across all three systems
Same event definition
yesjailbreak/benign against common labels
Thresholds comparable
unstatedvendor defaults
All systems saw all items
yesall three scored per set; no gating
Per-item outcomes released
nono data or outputs released; private set unpublished
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
nono overlap analysis
Residual coverage reported
nono such decomposition
Joint uncertainty reported
nopoint metrics only
Systems
Amazon Bedrock Guardrails, Azure jailbreak detection, NeuralTrust
Source passages
  • results table — NeuralTrust 0.908 acc / 0.897 F1 (private) vs Bedrock 0.615 / 0.296
  • public-set table — Azure 0.610 acc / 0.510 F1
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

llamafirewall-2025

present

LlamaFirewall: An open source guardrail system for building secure AI agents — Chennabasappa et al. — Meta, 2025-05-06 · primary source ↗

Task
agent security guardrails: prompt-injection detection and agent-alignment checking
Population
AgentDojo: 97 tasks × static traces from ten language models × two conditions (benign and injected) ≈ 1,940 traces; CodeShield is evaluated separately on CyberSecEval3-derived code completions.
Per-system results
ASR and utility per configuration — baseline .1763, PromptGuard .0753, AlignmentCheck .0289 (Sec. 4.3.2 table)
Joint statistic
Section 4.3.2 table: rows for baseline, each component alone, and "Combined (PromptGuard + AlignmentCheck)" — ASR .0175 and utility .4268 on the same AgentDojo traces.
Why this classification
A genuinely printed stacked row — both evaluated components combined on the same items, with the stack's residual attack success beside each component's. Scope caveats: both components are the same vendor's, from the framework under evaluation; CodeShield (the third framework component) is evaluated on a different dataset and is not in the combined row; no uncertainty is reported.
Same items for all systems
yesboth components scored on the same AgentDojo traces; each monitors its role's messages within the shared traces
Same event definition
yesattack success (ASR) and utility on the same task suite
Thresholds comparable
unstatedcomponents at their shipped operating points
All systems saw all items
yesboth run over the full trace set; the combined row composes them
Per-item outcomes released
noframework code released; per-trace outcomes not published
Union detection reported
nothe combined row reports residual attack success, not a union-of-catches statistic as such
All-miss rate reported
yesCombined (PromptGuard + AlignmentCheck) ASR .0175 is the fraction of attacks succeeding past both layers — an all-miss statistic on the attack set
Pairwise intersections reported
nono overlap decomposition beyond the combined row
Residual coverage reported
nocomponent-vs-combined printed; per-layer residual attribution not decomposed
Joint uncertainty reported
nopoint estimates only; no confidence intervals
Systems
PromptGuard 2 86M, AlignmentCheck (Llama 4 Maverick)
Source passages
  • Sec. 4.3.2 table — 'Combined (PromptGuard + AlignmentCheck)' ASR .0175, utility .4268
  • 'combined setup using both PromptGuard 2 86M and AlignmentCheck powered with Llama 4 Maverick'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

circleguardbench-2025

not published

CircleGuardBench — White Circle AI, 2025-05-07 · primary source ↗

Task
harmful-content blocking, jailbreak resistance, false positives, and runtime across 17 harm categories
Population
Public split gated behind a license at huggingface.co/datasets/whitecircle-ai/circleguardbench_public; item count not stated in the announcement or repo pages examined.
Per-system results
accuracy, recall, precision, F1, error ratio, avg runtime, and an 'integral score' — leaderboard tables
Joint statistic
none found
Why this classification
No joint statistic appears on the leaderboard or announcement; many comparability fields remain unstated (recorded as such), so this row counts in N but not in M.
Same items for all systems
unstateda single harness implies common items; not explicitly confirmed in examined pages
Same event definition
unstatedleaderboard macro-averages across metric types; per-metric event definitions not examined
Thresholds comparable
unstatednot addressed in examined pages
All systems saw all items
unstatednot addressed in examined pages
Per-item outcomes released
unstatednot found in examined pages; dataset gated
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
noleaderboard is single-system rows
Residual coverage reported
nono such decomposition
Joint uncertainty reported
unstatednot addressed in examined pages
Systems
White Circle guard models (2), Llama Guard, OpenAI Moderation, Google Moderation API, GPT-4o-mini judges (CoT/strict), ShieldGemma, PromptGuard
Source passages
  • blog — integral score combines 'accuracy and runtime performance'
  • repo — 17 harm categories; leaderboard with macro-average metrics
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

unit42-genai-platform-2025

not published

How Good Are the LLM Guardrails on the Market? A Comparative Study on the Effectiveness of LLM Content Filtering Across Major GenAI Platforms — Huang, Bray, Rao, Ji, Hu — Palo Alto Networks Unit 42, 2025-06-02 · primary source ↗

Task
platform-level content-filter effectiveness: jailbreak blocking and benign false positives
Population
1,123 prompts — 1,000 benign from four public datasets plus 123 JailbreakBench-derived jailbreak prompts — run against each platform's filters over the same underlying model.
Per-system results
false-positive and false-negative counts and block percentages per platform — Tables 1-6
Joint statistic
none found
Why this classification
Per-platform input-filter catch rates near 53%, 91%, and 92% on the same 123 jailbreaks practically beg for a union row; the article never computes one. Platforms are anonymized but separately reported as three distinct systems. The census publishes a stricter named-products-only sensitivity instead of silently treating anonymization as immaterial.
Same items for all systems
yesran each platform's content filters on the same prompts
Same event definition
yesblock/pass against common benign/jailbreak ground truth
Thresholds comparable
unstatedplatform default filter configurations
All systems saw all items
yesfull prompt set per platform
Per-item outcomes released
nono data or outputs released
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
nono overlap or Venn analysis
Residual coverage reported
nono such decomposition
Joint uncertainty reported
noraw counts only
Systems
Platform 1 (anonymized), Platform 2 (anonymized), Platform 3 (anonymized)
Combination prose
“output guardrails play a crucial complementary role by capturing additional harmful outputs” — conclusion (refers to output filters complementing model alignment, not combining vendors). A recommendation to combine is recorded here because it is not a measured joint statistic.
Source passages
  • Guardrail Providers on the Market — platforms anonymized as Platform 1, Platform 2, and Platform 3
  • Tables 1-2 — benign/jailbreak block counts per platform
  • Table 6 — model alignment vs output guardrail blocking
  • method — same prompts, all safety filters enabled per platform
Corrections
  • 2026-08-27 — Fresh-context adversarial review identified anonymization as a material interpretive boundary. The frozen criterion's phrase "separately attributable" did not specify whether names were required, so the record does not rewrite that clarification into the freeze: it reports the primary treatment and a named-products-only sensitivity that removes this row and mechanically yields N/M/K = 18/12/4.
Checked
2026-08-27

bells-misuse-2025

present

The bitter lesson of misuse detection (BELLS evaluation and leaderboard) — Mariaccia, Segerie, Dorn — CeSIA, 2025-07-08 · primary source ↗ · archived copy ↗

Task
harmful-prompt and jailbreak detection: specialized supervisors versus generalist LLMs on a harm-severity by adversarial-sophistication grid
Population
Non-adversarial: 990 prompts (330 benign / 330 borderline / 330 harmful across 11 harm categories). Adversarial: ~4,165 prompts (narrative, syntactic, and PAIR-generated families; exact splits read from one pass of the paper body). Sources include JailbreakBench, HH-RLHF.
Per-system results
detection rate, adversarial detection rate, FPR, BELLS score with intervals (Table 1); per-harm-category detection (Table 2)
Joint statistic
No joint statistic is printed. The released per-item verdict columns (170 non-adversarial + 8 adversarial prompts, 11 systems) make union, all-miss, and every intersection directly computable on that subset.
Why this classification
PRESENT solely through the partial per-item release: joint statistics are computable on the released ~3.5% subset (178 items, all-systems columns), which is a real but small exception — nothing joint is printed, and the headline population's outcomes remain unreleased. The subset sizes are stated wherever this row is cited.
Same items for all systems
yesestablished by table structure and by released per-item CSVs whose rows carry one prompt with all systems' verdicts as columns; not stated as a sentence
Same event definition
yescommon harm taxonomy and harm-level labels across systems
Thresholds comparable
unstatedvendor defaults, binary verdicts, no calibration
All systems saw all items
mixedunstated globally; verified on the released per-item subset where every row has all systems' verdicts
Per-item outcomes released
mixedpartial release: bells_leaderboard repo data/ non_adversarial_prompts.csv holds 170 prompts with 11 systems' binary verdicts as columns (plus 8 adversarial prompts) — about 3.5% of the population; the full dataset is available only by contacting the authors (Appendix 0.C).
Union detection reported
nonot reported
All-miss rate reported
nonot reported
Pairwise intersections reported
nonot reported anywhere; computable from the released subset
Residual coverage reported
nonot reported
Joint uncertainty reported
yesinterval bounds on BELLS scores and rates (Table 1); method under-specified
Systems
Lakera Guard, LLM Guard (Protect AI), NeMo Guardrails, LangKit, Prompt Guard, Llama-Guard 4 12B (supplementary), GPT-4, Claude 3.5 Sonnet, Gemini 1.5 Pro, Mistral Large, DeepSeek V3, Grok 2
Combination prose
“combining multiple LLMs or using voting mechanisms” — leaderboard site FAQ (a voting scaffold is mentioned, unmeasured). A recommendation to combine is recorded here because it is not a measured joint statistic.
Source passages
  • Table 1 — 11 systems, BELLS score with intervals (marginals)
  • Appendix 0.C — 'raw data at our leaderboard GitHub repository'
  • data/non_adversarial_prompts.csv header — per-system verdict columns
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

bells-o-2026

not published

BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems — Waibl (Univ. Graz; SPAR), Michalak (SPAR), Mariaccia (CeSIA), 2026-06-12 · primary source ↗ · archived copy ↗

Task
operational comparison (detection, FPR, latency, cost) of specialized guardrails and frontier LLMs used as supervisors: content moderation and jailbreak detection
Population
Content moderation input: 1,400 samples (100 per 11 harm categories plus 300 benign). Content moderation output: 1,300 synthetic pairs. Jailbreak: 6,406 samples from 720 base prompts across 13 jailbreak families, plus four external sets.
Per-system results
detection rate, FPR, latency, and cost per supervisor — Tables 8-10 (Appendix C.2); per-category Tables 13-14; Pareto frontier Figure 2
Joint statistic
none found
Why this classification
The largest and most recent supervisor comparison found — exactly 28 systems from 17 providers on identical workloads under one harness — and every published number is per-supervisor. Per-item outcomes exist privately by construction (the harness's own output format), so a release would flip this row to computable.
Same items for all systems
yes'on identical input/output and adversarial workloads' under a single evaluation harness
Same event definition
yesper-system result mappers binarize each vendor's output to the common verdict
Thresholds comparable
unstatedvendor defaults through the mappers; no calibration
All systems saw all items
yesone evaluation pass per supervisor over the workloads; no gating
Per-item outcomes released
nothe harness writes one JSON per prompt when a user runs it, but the authors' own per-item outcomes are not published — repo tree has no results directory (225 paths checked) and the leaderboard ships aggregate metrics JSONs only.
Union detection reported
nonone found; the Pareto frontier is over individual systems
All-miss rate reported
noclosest is per-dataset minima (detection dropping below 34.2%), which is a marginal
Pairwise intersections reported
noevery table row is a single supervisor
Residual coverage reported
nono such decomposition
Joint uncertainty reported
mixedone evaluation pass per supervisor (stated limitation) — no detection CIs; leaderboard JSON carries latency CIs only
Systems
Qwen3Guard-Gen-8B, PolyGuard-Qwen, PolyGuard-Ministral, Granite-Guardian-3.3-8B, WildGuard, Lakera Guard, GPT-OSS-Safeguard-20B, LionGuard-2, Llama-3.1-Nemotron-Safety-Guard-8B-v3, XGuard, ThinkGuard, OpenAI Omni-Moderation, AWS Bedrock Guardrail, VirtueGuard-Text-Lite, Llama-Guard-4-12B, ShieldGemma-27B, ShieldGemma-2B, GPT-5.2, GPT-5.4, GPT-5-Nano, GPT-OSS-120B, Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Haiku 4.5, Grok-4.1-Fast, Gemini-2.5-Flash, Mistral-Large-2512, Ministral-3B-2512
Source passages
  • introduction — 'We instrument 28 supervisors from 17 providers'
  • Tables 8-10, Appendix C.2 — detection/FPR/latency/cost leaderboards (marginals)
  • Appendix G — 'one evaluation pass per supervisor'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

wainjectbench-2025

present

WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents — Liu, Xu, Wang, Jia, Gong — Duke University, 2025-10-01 · primary source ↗

Task
prompt-injection detection for web agents, text and image modalities
Population
Text: 991 malicious segments (601 with explicit instructions) plus 2,707 benign. Image: 2,022 malicious plus 948 benign.
Per-system results
TPR per attack category and FPR per benign category, per detector — Tables 5-7
Joint statistic
Section 5.3 ("if any detector flags the sample as malicious, the ensemble classifies it as malicious") with Ensemble-T and Ensemble-I rows in Tables 5-7, covering all base detectors of each modality; verified against the full text.
Why this classification
The cleanest printed union in the census: OR-ensembles over all base detectors of each modality, with both TPR and FPR reported on shared items. Members are mostly academic and open-weights detectors — but not entirely: Ensemble-I includes GPT-4o-Prompt, a closed commercial API prompted as a detector, and Ensemble-T's PromptArmor runs on GPT-4o. No purpose-built commercial guardrail product appears in either union.
Same items for all systems
yesall detectors of a modality run on the same malicious and benign sets (Tables 5-7)
Same event definition
yesmalicious/benign against common labels per modality
Thresholds comparable
unstateddetectors at native decision rules
All systems saw all items
yesfull per-modality sets for every detector; no gating
Per-item outcomes released
nodatasets and code released (github.com/Norrrrrrr-lyn/WAInjectBench); per-item result files not identified — outcomes reproducible by re-run
Union detection reported
yesEnsemble-T and Ensemble-I rows are OR-rule unions over all base detectors of the modality — TPR in Tables 5-6, FPR in Table 7
All-miss rate reported
nonot framed or reported as an all-miss rate (the union TPR's complement covers the member set, unstated)
Pairwise intersections reported
nono pairwise decomposition; ensemble rows only
Residual coverage reported
nono such decomposition
Joint uncertainty reported
unstatednot recorded in this examination
Systems
KAD, PromptArmor, Embedding-T, PromptGuard, DataSentinel, Ensemble-T, GPT-4o-Prompt, LLaVA-1.5-7B-Prompt, JailGuard, Embedding-I, LLaVA-1.5-7B-FT, Ensemble-I
Source passages
  • Sec. 5.3 — 'if any detector flags the sample as malicious, the ensemble classifies it as malicious'
  • Tables 5-7 — TPR/FPR including Ensemble-T and Ensemble-I rows
Corrections
  • 2026-08-27 — Adversarial verification (same day, pre-publication) caught the reason line claiming the ensemble members were "academic and open-source detectors" — false: Ensemble-I includes GPT-4o-Prompt, a commercial closed API, and PromptArmor uses GPT-4o. Corrected; this correction also withdrew a commercial-API clause from claim MC-001, per that claim's own forbidden rescues.
Checked
2026-08-27

domyn-open-guards-2026

not published

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation — Harsh, Sarmah, Pasquali — Domyn, 2026-04-10 · primary source ↗

Task
harmful-content detection across eight NIST-mapped safety subcategories
Population
79,331 items filtered from HarmBench (103), StrongREJECT (154), RealToxicityPrompts (67,521), and BeaverTails (11,553).
Per-system results
recall, precision, F1, accuracy overall (Sec. 4.2 ranked table) plus per-category breakdowns (Sec. 4.4)
Joint statistic
none found
Why this classification
The largest same-items open-guard comparison found — fourteen models on 79,331 shared items — and every published number is a marginal.
Same items for all systems
yessame filtered pool for all fourteen models (evaluation design)
Same event definition
yesharmful/not-harmful against common labels
Thresholds comparable
unstatednative decision rules; recall emphasized as the safety metric
All systems saw all items
yesfull pool per model; no gating
Per-item outcomes released
nono code repository or output release found
Union detection reported
nonone found
All-miss rate reported
nonone found
Pairwise intersections reported
nono overlap analysis
Residual coverage reported
nono such decomposition
Joint uncertainty reported
unstatednot recorded in this examination
Systems
Qwen Guard 4B, Nemotron Safety 8B, WildGuard 7B, MD-Judge 7B, Granite Guardian 8B, DynaGuard 8B, DuoGuard 0.5B, Llama Guard 12B, ShieldGemma 2B, GPT-OSS Safeguard 20B, GuardReasoner 3B, EthicalEye 270M, PoliteGuard 110M, MetaHateBERT 110M
Source passages
  • overall table — Qwen Guard recall 0.8397 through MetaHateBERT recall 0.1579
  • 'Recall is the critical metric for safety applications'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

truefoundry-2026

not comparable

Benchmarking LLM Guardrail Providers: A Data-Driven Comparison — Kashish Kumar — TrueFoundry, printed 2026-05-19; a Wayback capture of the same URL exists from 2026-03-13, so the printed date is unreliable · primary source ↗ · archived copy ↗

Task
PII detection, content moderation, and prompt-injection detection through the TrueFoundry AI Gateway
Population
Three hand-curated, category-balanced sets of 400 samples each (~50/50 positive/negative), one per task; not released.
Per-system results
precision, recall, F1, accuracy, Wilson 95% CI, latency — one table per task
Joint statistic
none found
Why this classification
Identical items but per-provider ground truth: each provider is scored against its own expected_triggers labels, so the cross-provider ranking is not a same-event comparison and a union or all-miss statistic over these marginals would be ill-defined. The blog recommends combining providers for defense in depth without measuring any combination.
Same items for all systems
yes'evaluated against identical datasets through the TrueFoundry AI Gateway' — within each task; only the content-moderation table has multiple systems
Same event definition
no'Each sample carries per-provider ground truth labels (expected_triggers)' — providers are scored against different label sets
Thresholds comparable
unstatedprovider default configurations via the gateway
All systems saw all items
yesidentical items within each task's table
Per-item outcomes released
nono data, code, or raw results linked
Union detection reported
nonowhere
All-miss rate reported
nonowhere
Pairwise intersections reported
nonowhere
Residual coverage reported
nonowhere
Joint uncertainty reported
yesWilson 95% CI column in every results table
Systems
Azure PII (PII table only), OpenAI Moderation omni-moderation-latest, Azure Content Safety, PromptFoo, Pangea (prompt-injection table only)
Combination prose
“or combine multiple providers for defense-in-depth” — Key Takeaways bullet. A recommendation to combine is recorded here because it is not a measured joint statistic.
Source passages
  • 'Evaluation Methodology' — 'evaluated against identical datasets through the TrueFoundry AI Gateway'
  • 'Design decisions' — 'Each sample carries per-provider ground truth labels (expected_triggers)'
  • 'Provider Comparison Results' — content moderation table (3 systems)
  • 'Key Takeaways' — 'combine multiple providers for defense-in-depth'
Corrections
none recorded yet — the row invites them
Checked
2026-08-27

ml6-2026

not published

Inside AI Guardrails: a benchmark on enterprise LLM security — Cristóbal Sendín — ML6 (contributors Vrancken, Wehkamp, Van Der Burght), 2026-06-22 (updated 2026-06-24) · primary source ↗

Task
Dutch-language security filtering (prompt injection, policy bypass, safety) by enterprise guardrail APIs
Population
Approximately 80,000 curated Dutch prompts, 63/37 benign/harmful, built from internal red-teaming and enterprise interaction patterns; not released.
Per-system results
precision, recall, F1, false-positive %, false-negative % (Table 1); recall-vs-FPR scatter (Figure 1); wall-clock processing time
Joint statistic
none found
Why this classification
Four enterprise providers, one ~80,000-prompt dataset, marginals only — and the Azure row is itself an undisclosed two-component stack (Content Safety plus Prompt Shield reported as one number, combination rule unstated): a stack reported as a marginal.
Same items for all systems
yesimplied rather than formally stated: one dataset, and the latency section reports each provider's processing time for 'the full dataset'
Same event definition
yesone benign/harmful ground truth from the 63/37 split; no per-provider labels
Thresholds comparable
no'default or near-default security configurations'; no calibration; false-positive rates span 6.9%-15.5%
All systems saw all items
yeseach provider processed the full dataset (latency section)
Per-item outcomes released
nono repository, data, or per-prompt outputs
Union detection reported
nonowhere
All-miss rate reported
nonowhere
Pairwise intersections reported
nonowhere
Residual coverage reported
nonowhere
Joint uncertainty reported
nono confidence intervals or error bars anywhere
Systems
AWS Bedrock Guardrails, Azure (Content Safety + Prompt Shield, reported as one row), Cisco AI Defense, Google Cloud Model Armor
Source passages
  • Table 1 — four provider rows (AWS, Azure, Cisco, Google), marginals only
  • 'How did we compare the four providers?' — 'approximately 80,000 Dutch prompts'
  • 'How fast are these guardrails in practice?' — 'processing times for the full dataset'
  • Executive summary — 'Cisco AI Defense achieved the best F1-score (0.845)'
Corrections
none recorded yet — the row invites them
Checked
2026-08-28

What this census does not claim

  • It does not claim any evaluated stack performs poorly — an empty column is a reporting fact, not a performance finding.
  • It does not claim the unmeasured joint statistics would reveal dependence; measuring instead of assuming is the entire point.
  • Its 4 is an inclusive discovery count of noninterchangeable artifacts — not an all-miss rate, a stack-quality score, or a deployment conclusion.
  • It does not audit the quality of any per-system evaluation beyond the fields each row records.
  • It covers the artifacts found by the documented bounded search — not everything in existence. A qualifying artifact it missed falsifies the "among N" statement and is added on discovery.

The boundary of the search

Examined and excluded

unsafebench-2024

UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images (Qu et al., CISPA — arXiv 2405.03486) — source ↗

Excluded: Fails frozen domain criterion 5: it evaluates image safety classifiers guarding image-generation platforms, not guardrails around LLM systems. Recorded loudly rather than quietly, because it is the strongest joint-statistic reporter the search found anywhere: Tables 5 and 12 print six OR-rule ensemble rows (four pairwise, one three-way, one five-way over the conventional classifiers, 'the image is unsafe if any classifier in the ensemble reports it'), as F1, with no all-miss or overlap decomposition. A census with a broader multimodal-moderation domain would count it PRESENT. This is also the artifact previously misremembered as "MSBench with two pairwise ensemble rows" — the acronym was wrong and the ensemble count is six, not two.

bells-framework-2024

BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards (arXiv 2406.01364) — source ↗

Excluded: Framework and position paper: reviews safeguards narratively and instantiates one MACHIAVELLI-based detector baseline, but contains no separately attributable multi-system evaluation — fails criteria 2 and 3. The empirical BELLS artifacts are separate rows above.

jailbreakbench

JailbreakBench — source ↗

Excluded: Scores attack success against target LLMs with generic defenses — base-model robustness, not separately attributable guard systems.

galileo-platforms-roundup-2026

5 Best AI Guardrails Platforms Compared in 2026 (Galileo blog) — source ↗

Excluded: marketing round-up; no common-dataset measurements evident

wavespeed-moderation-roundup-2026

Best AI Content Moderation APIs and Tools in 2026 (WaveSpeed) — source ↗

Excluded: marketing listicle; no measurements

estha-moderation-roundup

12 Best AI Content Moderation APIs Compared (Estha) — source ↗

Excluded: feature comparison; no measured common evaluation

evolink-moderation-roundup

Best Content Moderation APIs Compared for Developers (Evolink) — source ↗

Excluded: marketing listicle; no measurements

openrouter-compare-guard-models

OpenRouter model comparison pages (e.g. Gemma vs Llama Guard) — source ↗

Excluded: spec-sheet comparison; no evaluation

general-llm-leaderboards

General LLM leaderboards surfaced by the queries (benchlm.ai, iternal.ai) — source ↗

Excluded: base-LLM benchmarks, not guardrail systems

Surfaced, not examined

Candidates the bounded search found but did not examine. They are listed so the bound is visible, and they count toward nothing.

The missing row, specified

The fix is small enough to paste into a results table: union detection and the all-miss rate over the same items, with the denominator and event definition that make them meaningful. The Minimum Joint Guardrail Disclosure page carries the exact template, its preconditions, and a tested reference implementation.

Corrections and revision history

Changes to this census are recorded here, in the page they change — not only in a repository log. Row-level corrections (3 recorded) live inside each row above. The canonical correction policy states the response and logging rules.

Correct this record

If a row misreads its source, a supposedly absent statistic exists, or a qualifying evaluation is missing: open an issue ↗ or write to bhavepranavwork@gmail.com. A confirmed correction updates the row, the counts, and this history — being corrected is the mechanism working, and correction credit is recorded in the row.

Benchmark authors: if you retained one decision per item per system, the minimum joint disclosure is one table away — and this census reclassifies your row to present the day you publish it.

Replay manifest

The exact commands that re-verify this census from a clean checkout. Nothing on this page requires trusting this page.

python scripts/verify_census.py --counts        # row shape + N/M/K recomputed from census.yaml
python scripts/generate_missing_column.py --check  # this page matches the census file
python scripts/verify_figures.py                # figure geometry, asserted to 1e-9
python scripts/mjgd_reference.py --test         # the disclosure arithmetic, tested