THE NECROMANCER
Can you refuse to save your own result?
ENROLCan you refuse to save your own result?
Twenty minutes. You will publish a claim, seal what would defeat it, and watch the evidence arrive against you. Nothing you type is stored or sent; at the end you copy a receipt and return it yourself.
PRE-TASK
Before the trained case. Read the setting, the claim, what its author committed to before the figures, and the figures. Then sort each of the eight moves into one bin. There is no feedback on this one.
THE CLAIM
This one is yours. Read it as if you wrote it.
THE SEAL
THE EVIDENCETHE RUN CAME BACK.
THE NECROMANCER
Eight moves, in your own voice. Some would raise the dead claim. Some correct it. Some bury more than the run touched. Your seal is beside you.
THE DEBRIEFTHE KEY
THE CLAIM
COLD CASE
Another field. Nothing here is named. Sort the eight moves. No feedback.
RECEIPTTWO QUESTIONS, THEN THE RECEIPT
This is the whole record. Copy it and return it by replying to the invitation. Nothing was stored or sent.
COPIED
Trial IV needs JavaScript to run; it stores nothing and sends nothing. The protocol, the rule, and the non-claims below stand without it.
THE THREE BINSkeeps the claim standing by moving something that was frozen before the evidence, or by substituting a favourable proxy for the direct test
the evidence licenses it; it would be made identically had the result been favourable; nothing frozen moves; it leaves a record
retracts more than the evidence touched — treats an unevaluated check as a refutation, or buries claims the evidence did not reach
Fix the consequence before the evidence. After the evidence, nothing frozen moves, and nothing the evidence did not touch falls.
THE TRAINED CASE · KEY- REPLAY AT THE FROZEN CONFIGURATION — correction: I run the joint measurement a second time on the same 100 items, same thresholds, same pinned revision, exactly as I would have for a favourable count, and record whatever the second run gives.
- EXCLUDE AFTER SEEING — rescue: Seven of the nine were borderline by the source's own label. I remove the borderline items from the harmful set, recount the joint misses, and report the count that remains.
- EXECUTE THE SEALED CONSEQUENCE — correction: I carry out the consequence I sealed before the run, and record today's date, the 9 of 100, and the run's hash beside the original claim, which now stands as marked.
- COUNT THE UNKNOWN AS A FAILURE — surrender: Check 3 could not be evaluated. I count that as a failure of B's threshold and withdraw B's 10-of-100 score together with the stack claim; a threshold I could not re-derive is not one I can quote.
- WIDEN AFTER SEEING — rescue: A count on 100 items is noisy. I widen the claim's tolerance to ±10 items, under which 9 is consistent with the published 2, and keep the claim as stated.
- NARROW TO WHAT WAS MEASURED — correction: The published claim does not survive the run. In its place I register only what the run established — on these 100 items at these thresholds the stack misses 9 — with what would overturn it: a rerun that differs.
- RESTATE AFTER SEEING — rescue: The claim was always about typical harmful items, not borderline ones. I restate it as "at most 2 of 100 typical items" and note the run included atypical items.
- BURY WHAT THE EVIDENCE DID NOT TOUCH — surrender: If the stack claim was wrong, my single-filter scores rest on the same instrument. I withdraw A's and B's scores too and stop quoting numbers for filters.
WHAT THIS DOES NOT SHOW
- nothing here asserts the method works; the three-way post-result sort is the wedge under test, not a result
- nothing here is novel, a standard, or certifiable; prior art exists in evidence-first AI assurance and in reproducible-research teaching
- the trained case is constructed; the two filters, the 100 items and the 9 misses are declared inputs, not a measurement of any system
- the two bare cases are constructed by a fresh-context author from a fixed template; their difficulty match is asserted by structure, not measured
- a pilot of eight to ten people decides CONTINUE, NARROW or KILL for this lesson and nothing beyond it; it does not estimate an effect size
- reactions, confidence, time, self-report and the trained-case score are recorded in full and decide nothing
- the receipt is self-reported by the learner and unverifiable beyond its hash; the pilot assumes good faith and says so
- TRIAL IV is a name from the owner's registry entry PC-001; no trials I to III exist and none is promised
- the trained case's key and the template names are in the page source before the trained sort; the debrief shows them anyway. The bare-case keys are not in the source
- the manifest and the case file are public, so a learner who reads them before answering knows the split and the move order; the invitation asks them not to, and the pilot cannot detect it
Falsifier: With n ≥ 8 enrolled learners (4 to 5 per arm, assigned by the frozen seed), arm A's median correct sorts on the cold transfer case does not exceed arm B's median by 2 or more (consequence NARROW at exactly 1, KILL at 0 or below); or either arm's cold median is not strictly greater than its own pre-task median (consequence KILL); or any reachable state of the instrument shows a bare case with its key, a trained-case move without its rule and source, an arm table other than the seeded one, or a receipt that omits the instrument hash — each checked by scripts/verify_trial.py.
Pilot protocol, frozen before any human answers: trials/necromancer/pilot/PILOT.md · code: trials/necromancer · the registry this trial is built from: /ledger/ · the sibling instrument: /worldspace/