Evidence & reproducibility

What is demonstrated, what is internally reproducible, and what remains open.

The goal is to make the evidence boundary easier to inspect than the narrative.

01 — Ledger

Evidence ledger.

internally reproducible is not independently reproduced. Retained negative results are listed alongside positive ones.

QuestionEvidenceReproducibility statusBoundary
R.01
Can information acquisition be premature?
Frozen 15-state diagnostic; full REC 15/15; principal timing/consequence ablations collapse pre-boundary performance.Offline paper-scoped artifact with raw outputs, hashes, evaluator, reference implementation, and verification scripts.Independent external reproduction and broad real-world prevalence are not yet established.
R.02
Can passive value-of-information choose the wrong RFQ policy?
Formal representability boundary; 10,080 heterogeneous tests; fresh 5,000-world IID holdout; mechanism controls and shifts.Exact enumeration; separated discovery/holdout seeds; frozen holdout ranges.Stylized mechanism test, not an empirical estimate of real RFQ leakage frequency.
R.03
When does private-credit model error become decision-critical?
Frozen 20-world holdout; line-cut control; 14/20 failures under omitted working-capital mechanism; stressed optimum feasible 20/20.Pre-specified holdout plus separately frozen, source-anchored 40-world follow-up.Original/follow-up are not hash-continuous; synthetic environment; no deployed-system prevalence claim.
Negative result
Is REC a new generic planning primitive?
Negative result: exact depth-3 expectimax solves 27/27 small worlds.Retained explicitly in the reproduction package.No generic-planner novelty claim.
AEI
Does the broader AEI thesis generalize?
Multiple internal hidden-rule, rare-event, constraint, document, and adaptive-environment evaluations.Large private forensic archive with frozen specs, hashes, raw outputs, and reconstruction material.External replication, externally specified worlds, and real-world validation remain decisive next tests.
02 — Research standards

How claims are handled.

S.1

Freeze before confirmatory evaluation

Where a result is treated as prospective evidence, the environment, seed rule, analysis, or evaluation criteria are fixed before the corresponding holdout is run.

S.2

Keep negative evidence

Generic-planner successes, oracle contamination, non-identifiability, failed controls, and null results are retained rather than silently dropped.

S.3

Separate discovery from confirmation

Development worlds and post-hoc follow-ups are not relabeled as prospective confirmation.

S.4

State external-validity limits

Synthetic mechanism tests are not presented as prevalence estimates for deployed systems.

03 — Artifacts

Public, paper-scoped, and private layers.

Layer 1

Public

Research summaries, selected evidence, neutral manuscript metadata, and safe research notes.

Layer 2

Paper-scoped

REC paper-scoped offline artifact; finance reproduction packages can be published as separate safe repositories.

Layer 3 · not disclosed

Private

Full AEI architecture, reconstruction-grade implementation details, complete internal evaluation suite, and successor design.