Research note

Why Prediction Is Not Enough for Decisions Under Hidden Rules

A system can be excellent at predicting outcomes inside the wrong model of the world. In consequential decisions, that distinction matters.

I started from a practical problem: build agents that make economic decisions rather than merely describe them. The first surprise was that better model reasoning did not reliably remove the failures I cared about. In some environments, a system could explain the relevant rule once it was made explicit and still fail when it had to determine which rule governed the decision before acting.

That pushed the research away from a simple “make the model smarter” framing. The recurring issue looked more structural. A decision can depend on at least four different questions: what the current state is, which mechanism governs its evolution, whether the available evidence supports that mechanism in the region the plan will visit, and what happens in rare or adversarial states that average-case prediction barely sees.

Prediction and identification are different operations

Suppose two economic mechanisms generate similar observations over the data seen so far but imply different feasible actions once a covenant, cash threshold, or rare transition becomes active. A predictor can fit the observed outcomes well without resolving the distinction that actually controls the decision. More samples help only if they contain information that separates the mechanisms.

Fig. — Prediction vs. identificationSchematic
Schematic: two mechanisms agree on observed data and diverge in the decision region Data points follow a single curve in the observed region. Beyond it, mechanism A rises while mechanism B falls through a feasibility boundary. OBSERVED SO FARWHERE THE ACTION IS TAKEN FEASIBILITY BOUNDARY Mechanism A Mechanism B action infeasible under B STATE / EXPOSURE →
SchematicTwo mechanisms agree on the data seen so far and diverge where the action is taken. Illustrative, not experimental data.

This is why some controlled hidden-rule experiments remained difficult for outcome learners even when the rule-conditioned action itself was easy. Once the governing rule was revealed, the comparison model often knew what to do. The failure occurred earlier: identifying which structure should govern the action.

Information acquisition can change the world

The same distinction appears in a different form when an agent seeks more information. A standard abstraction treats sensing as something that improves the belief state, perhaps at a cost. But in many economic systems an inquiry is an intervention. Requesting a quote can reveal intent. Asking a counterparty a question can change behavior. A diagnostic action can consume a resource or remove a later option.

Then “more information” is not a free refinement of the same decision problem. The acquisition action can create a new decision problem. That motivates the work on premature epistemic action and endogenous value of information.

Constraint geometry matters more than average model error

Private-credit stress tests produced another version of the same lesson. A model can be imperfect in many ways without changing the decision. The error becomes critical when it moves a boundary the chosen policy actually reaches. The interesting quantity is therefore not generic model error but whether the error changes the feasible set along the policy trajectory.

Rare events create a related evaluation problem

Average-case performance can hide the regimes that dominate economic value. As catastrophic transitions become rarer, an outcome learner may see less direct evidence of the event even while the structural consequence remains decisive.

What I currently believe

I do not think the evidence justifies a sweeping claim that prediction is unimportant or that one architecture solves decision-making generally. The narrower thesis is that some consequential decision problems require explicit work above prediction: identifying governing structure, establishing support, stress-testing relevant regimes, planning under the resulting model, and verifying execution against the authoritative environment.

That thesis connects the current research lines. Premature epistemic action studies when an information-seeking action should be delayed. The RFQ work studies when passive value-of-information is the wrong abstraction because inquiry changes downstream values. The private-credit work studies when representation error becomes consequential because it moves a reached feasibility boundary. The broader AEI program asks whether these recurring patterns can be turned into a more reliable decision architecture.

What would change my mind

The strongest next evidence would not be another hand-designed success. It would be an externally specified evaluation where the governing mechanisms, world families, and scoring are fixed independently, then tested across model providers and with enough real or semi-real structure to make overfitting difficult. If a simpler predictive or generic planning baseline matches the structural methods under those conditions, the broader architectural thesis should shrink accordingly.