Engagements
Guardrail evaluation design
- For
- Benchmark authors, evaluation teams, and platform safety groups about to publish or refresh a guardrail comparison.
- Problem
- Per-detector scores are collected on different items, at unmatched operating points, and the stack's own behaviour is never measured — so the published table cannot answer the question a deployment actually asks.
- You get
- An evaluation design: shared item set, one event definition, matched operating points, declared exposure conditions, and the measurement plan for union detection, all-miss rate, residual coverage, and uncertainty — plus the results table shaped so the joint row is a first-class output rather than an afterthought.
- Boundary
- A design and a measurement plan. Not a certification, not a compliance sign-off, and not a claim that any system or stack is safe. Running the evaluation and interpreting it stay yours.
- Next
- Send the results table you already publish.
AI claim and evidence audit
- For
- Teams whose safety, accuracy, or robustness claim is about to face procurement, regulators, press, or a paper reviewer.
- Problem
- The claim is defensible and the reasoning behind it lives in someone's head. Nobody can say precisely what it assumes, what would refute it, or which weakening moves would be cheating.
- You get
- One claim taken apart: an evidence map, the assumptions it rests on, a reproduction attempt from the bound sources, an explicit falsifier with its consequence, the rescues ruled out in advance, the non-claims, and a decision-facing summary a non-specialist can act on.
- Boundary
- An audit of what your evidence establishes, not an endorsement that it is true. If the evidence supports less than the claim says, the deliverable says so — that is what you are buying.
- Next
- Name one claim you would least like to be wrong about.
Receipt and provenance threat model
- For
- Teams shipping attestations, content credentials, model cards, audit logs, or governance receipts.
- Problem
- A receipt proves something narrower than the thing people will read it as proving, and the gap is only discovered when someone relies on it.
- You get
- A written boundary: exactly what the receipt identifies, which transformations it survives, which leave it ambiguous, what a verifier can and cannot conclude from a pass, and the failure modes that look like successes.
- Boundary
- A threat model for what the artifact establishes. Not a security certification, not a penetration test, and not an assurance that the system is sound.
- Next
- Send the receipt format and one verifier you rely on.
What none of this is
No engagement here certifies a system, approves it for compliance, or asserts that a model, guardrail, or stack is safe. Every deliverable states its own boundary in writing, and every empirical figure it contains carries the scope of the evidence behind it. If you need a certificate, I am the wrong person; if you need to know what your evidence actually supports, that is the entire offer.
The public record is the sample of the work: a bounded census with a live correction route, a claim ledger where every proposition carries a falsifier, and a disclosure schema with a tested reference implementation.
Start a conversation
Email is the fastest route. A concrete artifact — a results table, a claim, a receipt format — gets a more useful reply than a general enquiry.