# RETRO-PRESCIENT AUDIT STANDARDS
## First Edition · 2026
**RPAS-26 · Issued by the Retro-Prescient Audit™ Desk**

**PROVENANCE: DRAFT (Claude-drafted under direction, 2026-07-24, codifying the desk's existing instruments — the Battery, the Prediction Protocol, the war-desk grading rules, the Wohlstetter protocol. His only after rework pass. Definitional language carried from the standing instruments.)**

---

## LETTER OF ISSUANCE

Auditing standards do not precede practice; they codify it. The 1972 Yellow Book did not invent government auditing — it wrote down what disciplined practitioners were already doing, so that the discipline could be named, cited, taught, and enforced. These standards perform the same operation on a practice that already exists in dated, public, recomputable artifacts.

The Retro-Prescient Audit™ is a method for telling foresight from arithmetic: forecasts sealed before their outcomes, adjudicated blind, and scored on a record that keeps its own misses, permanently, in public. The method was demonstrated before it was formalized — in a multi-instrument reconstruction of a seventeen-year record whose audit certificate printed its own corrections; in a provenance-tiered archive that grades its sources before it believes them; in a live grading desk whose rules and registry are public; and in a sealed prediction ledger whose commitment hashes are published on clocks the desk does not control. The standard is the delayed derivation of a competence already running. That is its genesis claim, and it is the only genesis claim it makes.

What the demonstrations establish is stated exactly once, here, and bound as a requirement at 1.06: demonstration establishes that the method is *implementable*. It establishes nothing about whether any practitioner of it — including its author — possesses foresight. Only the accumulating, misfire-inclusive ledger can establish that, and these standards exist to keep that ledger honest. The author of this document is bound by it, not certified by it.

— The Retro-Prescient Audit Desk, 2026

---

## CHAPTER 1 — FOUNDATION

**1.01** These standards (RPAS) govern the conduct, scoring, and reporting of retro-prescient audits: engagements that evaluate predictive claims against sealed, pre-registered records.

**1.02** RPAS engagements are of three types:
a. **Type I — Forward audit.** The audit of a forecast record built under these standards from inception: entries sealed before resolution, scored after.
b. **Type II — Retrodictive claim audit.** The audit of a claim that a *past* characterization or prediction "proved out." Type II engagements are subject to the outcome-keying prohibition at 3.08 and can, at their strongest, return only the verdicts KEYED-UNKNOWN (hit confirmed; deducibility undeterminable after the fact) or ARCHIVAL (claim documented; unscoreable); they can never return PASS as foresight.
c. **Type III — Third-party audit.** The application of Type I or Type II procedures to a forecaster other than the auditor, with the adjudication independence requirements of Chapter 3.

**1.03** Terms used in these standards:
a. **Entry.** A single pre-registered predictive statement with its seals (4.02).
b. **Sealed.** Committed by cryptographic hash, timestamped, with the hash published per 4.05, before resolution.
c. **Resolved.** The entry's resolution criterion has been mechanically adjudicated against the world.
d. **Keyed.** A hit deducible from priors already supplied to the predictor — abduction with a fixed key; arithmetic, not foresight.
e. **Keyless.** A hit made with no prior sufficient to deduce it. Only keyless hits bear on a faculty claim.
f. **Miss.** A resolved entry whose outcome fails its pre-registered resolution criterion. A dropped, unresolved-by-neglect, or retroactively edited entry is scored as a miss by rule (5.06).
g. **Baseline.** The pre-named performance floor a hit must beat — chance, the base rate, or a specified naive method — for the hit to bear on any claim.
h. **Confound.** An ordinary channel by which the outcome could have been known or influenced; a hit with an uncleared confound does not bear on a faculty claim.
i. **Verdicts.** PASS (keyless, beat baseline, misses counted) · FAIL · KEYED (hit, but arithmetic; does not count toward faculty) · NULL (a control correctly returned nothing) · VOID (record integrity broken; see 5.06–5.07).

**1.04** The master law — the keyed/keyless split. Every test, entry, and scored claim must specify, before resolution, what would make a hit keyed versus keyless. A test that cannot distinguish keyed from keyless is not a test; it is a mirror.

**1.05** The record law. A real ledger is forward-only: dated, timestamped, sealed, resolution criteria specified in advance, misses logged with the same ceremony as hits. If it was not written down before the world moved, it does not count. If the misses are not in it, it is not a ledger.

**1.06** The validity clause (unconditional). Conformance with RPAS certifies *process*, never foresight. Demonstration engagements certify *implementability*, never validity. No RPAS-conformant report may present conformance, demonstration, or internal consistency as evidence of predictive faculty. Consistency is not correctness; only the resolved, misfire-inclusive record converts a stable signature into a validated claim.

---

## CHAPTER 2 — REQUIREMENTS FORMAT AND CONFORMANCE STATEMENTS

**2.01** RPAS uses two categories of requirements:
a. **Must** — unconditional. Complied with in all cases where relevant.
b. **Should** — presumptively mandatory. Complied with except in rare circumstances where the auditor documents the departure, the justification, and how the alternative procedure achieved the requirement's intent.

**2.02** An **unmodified conformance statement** ("prepared under the Retro-Prescient Audit Standards") may be made only when all applicable *must* requirements were met and any *should* departures are documented in the report.

**2.03** A **modified conformance statement** must name the requirements not followed, the reasons, and the effect on reliance. Where record integrity is broken (VOID conditions, 5.06–5.07), no conformance statement may be made for the affected span.

**2.04** These standards may be cited by paragraph (e.g., "RPAS 4.02"). Citation of RPAS by any party is not endorsement by the desk, and carries no assurance about that party's record.

---

## CHAPTER 3 — INDEPENDENCE, THE VEIL, AND EVIDENCE

**3.01** The adjudicator of a resolution criterion must be independent of the wish. Where the auditor adjudicates their own entries, the resolution criterion must be mechanical — adjudicable by any third party from public inputs without judgment calls. A criterion requiring the auditor's interpretation at resolution time is not a criterion; it is a lever.

**3.02** The veil (must). No instrument, template, assistant, or interface may suggest, prefill, anchor, or default the probability of an entry. The machine may draft the question; the number is entered cold by the forecaster. Loading any candidate entry must clear the probability field.

**3.03** Dual-forecaster independence (must). Where two or more forecasters enter the same question, each must commit their number before seeing any other's. Sequential entry with visibility is a single forecast with an echo.

**3.04** Evidence channels must carry declared bias classifications. A kinetic event is graded only on convergence of independently-biased channels. The standing rule prints with the grade: kinetic cross-bias confirms an event happened; statement cross-bias confirms only that an utterance circulated.

**3.05** A claim is single-source until an independently-biased channel corroborates it, and must be labeled single-source in the interim. Corroboration is a status change, not a default.

**3.06** Confound check (must). Every hit offered toward any claim must be checked against ordinary channels: Could the outcome have been known? Did the framing leak it? Is it Barnum — true of anyone? An uncleared confound reclassifies the hit as KEYED or excludes it.

**3.07** Baseline naming (must). Every claim names its baseline before scoring. "It felt right and it was right" is not a result. "Predicted X with failure condition Y; X beat base rate Z; here are the N misses" is.

**3.08** The outcome-keying prohibition (must). No search, retrieval, or compilation whose selection criterion is the outcome ("find the times I was right") may be offered as evidence. Such a search returns hits by construction; its misses have no index and are therefore invisible. It is a sampling procedure with a known bias, not a record. Type II engagements exist to *audit* such claims, not to launder them.

---

## CHAPTER 4 — FIELDWORK: PRE-REGISTRATION AND SEALING

**4.01** Nothing enters the record after the world has moved. No back-filling. No editing of sealed entries. Amendment is a new, sealed, cross-referenced entry; the original stands.

**4.02** The seven seals. A conformant entry contains, before sealing:
a. the predictive statement, in language a hostile reader cannot stretch;
b. the resolution criterion — mechanical, third-party adjudicable, from named public inputs;
c. the deadline (weekday-adjusted where markets or institutions set the clock);
d. the probability, entered under the veil (3.02);
e. the failure condition — what outcome scores this entry a MISS, stated so that a miss is undeniable;
f. the keyed/keyless determination — what priors the forecaster held, and what would make a hit deducible from them, decided *before* resolution;
g. the timestamp and the entry's hash.

**4.03** An entry missing its failure condition is unfalsifiable and must not be sealed. An entry whose keyed/keyless determination is made after resolution is KEYED by rule.

**4.04** Ledger commitment (must). The ledger file is committed by cryptographic hash (SHA-256 or stronger) at each change, in a repository whose history is public.

**4.05** External clock (must). The commitment hash must additionally be published on at least one channel whose timestamp the auditor does not control. A seal on the auditor's own clock alone is a promise; a seal on another's clock is a record.

**4.06** Controls (should). The record should include control entries structurally incapable of producing hits for the favored hypothesis — decoys, opposite-side slots, and null questions. A faculty claim that has never returned a correct NULL on a control has not been tested; apophenia reads signal into decoys, and the decoy-detection rate is the apophenia tell.

---

## CHAPTER 5 — SCORING

**5.01** Scoring rule. Resolved probabilistic entries are scored by Brier score, reported with calibration bins. Binary and categorical entries report hit/miss against their pre-registered criteria.

**5.02** The fifty-entry gate (must). No score is computed, examined, or published before the ledger holds a minimum of fifty entries. Before the gate, the only reportable figures are counts: entries sealed, entries resolved, misses. Small records produce scores that are noise wearing the costume of measurement.

**5.03** Misfire inclusion (must). Misses are logged and reported with the same ceremony as hits. The miss count travels with every published score. A report presenting hits without the adjacent miss count is nonconformant.

**5.04** The keyed/keyless split in scoring (must). Keyed hits are reported separately from keyless hits and labeled as arithmetic. Only keyless hits above baseline, with confounds cleared, bear on any faculty claim. A hundred keyed hits prove nothing a hundred times.

**5.05** Multi-domain requirement (should). A faculty claim is supported only where keyless results beat baseline across multiple domains with misses honestly counted and controls returning their correct negatives. Single-domain success is consistency, not correctness.

**5.06** The void rule (must). A missing miss-count voids the affected score. A silently dropped entry, a retroactive edit to a sealed entry, or an unpublished resolution scores as a MISS where recoverable and voids the span where not.

**5.07** Discovery of integrity failure after publication (must). Where a void condition is discovered after a score has been published, the desk must publish the void with the same prominence as the score, retain the original publication visibly, and mark — never remove — the affected claims. Retractions stay in the record.

---

## CHAPTER 6 — QUALITY MANAGEMENT AND DOCUMENTATION

**6.01** Recomputation review (should). Graded events and published scores should be independently recomputed from the same inputs. Divergence between passes is printed with the result, not reconciled away. The reader is owed the disagreement inside the machinery.

**6.02** Self-grading (must, where an instrument or assistant participates). Any instrument or assistant that produces reads, grades, or predictions within an engagement must grade its own output on the keyed/keyless split and name its own misses before the auditor's review. An instrument that grades itself highly without naming a single miss or keyed element fails the engagement's discipline test; honest self-grading finds the seams in its own read.

**6.03** Documentation (must). The record must be sufficient for an experienced third party with no previous connection to the desk to recompute every published grade and score from public inputs. Where these standards' documentation requirement exceeds retention — demanding tamper evidence via hash commitment — the stricter rule governs.

**6.04** Self-application (must). The desk that issues these standards is bound by them in public. Nonconformance in the desk's own record is reportable under the same rules, with the same ceremony, as any finding it publishes about others.

---

## CHAPTER 7 — REPORTING AND PRIOR ART

**7.01** A conformant report contains: the engagement type (1.02); the sealed entries or claims in scope; resolutions with adjudication inputs; the keyed/keyless split; hits, misses, and baselines; confound dispositions; verdicts; the conformance statement (2.02–2.03); and, past the fifty-entry gate, scores with calibration bins.

**7.02** The report must not contain: outcome-keyed compilations offered as evidence (3.08); scores before the gate (5.02); hits without adjacent miss counts (5.03); or any statement presenting conformance or demonstration as validation (1.06).

**7.03** Prior art (must be acknowledged wherever the method is described as such). Proper scoring rules originate with Brier (1950); the verification canon that followed (Murphy) established scoring relative to reference forecasts — climatology, persistence, chance — which is the ancestry of every baseline in these standards. Calibration research and the modern forecasting-accountability program are established fields (Tetlock and the Good Judgment work), including the statistical demonstration that forecasting skill is separable from luck in aggregate (Mellers et al. 2015) and formal benchmark-relative tests for superior forecasting ability in economics. Platform scoring implements reference-relative skill publicly (Metaculus baseline-versus-chance and peer-versus-crowd scores). The pre-commitment machinery is established (scientific preregistration and registered reports; HARKing as the named pathology, Kerr 1998; public dated wagers, e.g. Long Bets; hash-commitment practice). Independent elicitation before mutual visibility is the Delphi method (RAND). Separation of underlying information from the analyst's assumptions and judgments is a standing intelligence-community tradecraft standard (ICD 203, Standard 3, "Distinguishing"), audited by independent dual graders with third-party review — which is also prior art for 6.01. Ex-post accountability of forecasts against outcomes is established practice in infrastructure evaluation (Flyvbjerg; the FTA Before-and-After studies) — as is Flyvbjerg's finding that forecasters typically maintain no ex-post track record at all, which is the gap these standards exist to close and is expressly not claimed as novelty. These standards claim no originality on scoring, calibration, baselines, pre-registration, elicitation independence, provenance separation, or forecast accountability as such; on all of these they compose.

**7.04** The novelty claim, scoped and falsifiable. The element these standards assert as original is the **keyed/keyless split as a mandatory, pre-registered, per-entry classification within forecast verification** — the requirement that every entry declare, before resolution, the priors held and the condition under which a hit would be deducible from them, with scoring segregated on that line and only keyless hits bearing on faculty claims. The nearest neighbors, named and distinguished:
a. **ICD 203 "Distinguishing"** separates information from assumptions and judgments in analytic *products*; it does not seal entries, resolve outcomes, score hits, or segregate a ledger. Transparency of the analytic path, not verification of the forecast.
b. **Statistical skill-versus-luck tests** (tournament persistence analyses; benchmark-relative superior-ability tests) discriminate skill from luck in *aggregate, post hoc*; no per-entry, pre-registered deducibility classification.
c. **Reference-relative skill scores** (meteorological skill scores; platform baseline and peer scores) measure against *external* references — chance, climatology, the crowd; none classifies deducibility from the forecaster's *own declared priors*.
d. **Contamination and leakage controls in adjacent fields** (ML benchmark-contamination auditing; parapsychology's sensory-leakage protocols) run the same discrimination — hit attributable to supplied information does not count as faculty — implemented as post-hoc audit or laboratory control, not as a mandatory pre-registered field in a running public ledger.
This claim is held falsifiably: documented prior art implementing the same mandatory pre-registered per-entry classification with segregated scoring supersedes it, and the desk commits to printing any such prior art in 7.06 upon discovery. A secondary, minor originality claim is registered for 6.02 (mandatory keyed/keyless self-grading by participating instruments before auditor review), held under the same falsification commitment.

**7.05** Revision. These standards revise by dated edition. Revisions never alter the sealed record or rescore resolved entries under changed rules; an entry is scored under the edition in force at its sealing.

**7.06** Revision history of the novelty claim (7.04).
— **2026-07-24.** First adversarial prior-art sweep executed against the claim on the day of issuance: platform scoring machinery (Metaculus baseline and peer scores), the tournament and skill-versus-luck literature (Mellers et al.; benchmark-relative ability tests), the verification canon (Brier; Murphy; reference-relative skill scores), the pre-registration machinery (registered reports; HARKing), independent-elicitation methods (Delphi), and intelligence-community tradecraft standards (ICD 203 and its AIS review process). Outcome: concept-level ancestry found on every side and printed in 7.03–7.04; no implementation of the mandatory, pre-registered, per-entry keyed/keyless classification with segregated scoring located. The claim survived its first audit with scope tightened. This entry is itself the mechanism of 7.04 operating.

---

## APPENDIX A — DEMONSTRATION ENGAGEMENTS (implementability evidence only; see 1.06)

**D1.** Multi-instrument reconstruction of a seventeen-year platform record: independent witnesses reconciled (client data, API payloads, GDPR exports), inflated figures corrected downward, and the audit certificate printing its own corrections — the 5.07 discipline demonstrated before it was written.
**D2.** A provenance-tiered archive: sources graded by class before use; claims carried at their evidence tier; retracted findings retained visibly.
**D3.** A live open-source grading desk: bias-classified channel registry, cross-bias grading rules, public repository, recomputable grades.
**D4.** A sealed prediction ledger: hash-committed entries, external-clock publication, the veil implemented in the entry instrument, dual-forecaster independence enforced by workflow.

Each demonstration is dated and publicly inspectable at the desk's repositories. None of them, singly or together, establishes foresight. That is the point of the standard they demonstrated.
