Scientific forecast evaluation

Does human context improve the model—and does the market add more?

Two frozen arms answer that question on later event families: human + model, then the same system + a matched prediction-market sensor. Every input lane stays separately scored.

not eligible no prospective matched human outcomes

Prospective resolved human forecasts

0

Only blind submissions made before the comparison lanes were visible can enter.

Already-public receipts excluded

7

They cannot be backfilled as human estimates without anchoring and hindsight contamination.

Production hybrid probability

None

No weight or combined number is published before both held-out studies pass.

Current scientific result

7 current receipts were already public before the human lane opened and are therefore excluded. No prospective resolved human forecast is available; both hybrid arms, every learned weight, and the combined probability remain unavailable.

MacroGuru has preregistered the comparison, not demonstrated hybrid superiority. A negative result will remain public.

The two-arm test

Without prediction market

Blind analyst point probability + statistical-model point probability. This tests whether first-party context improves the model on its own.

With prediction market

The same two lanes + a pre-lock matched market point. Market data is leverage, but cannot hide whether MacroGuru adds independent value.

Preregistered temporal holdouts

WindowHeld-out periodEarlier training familiesHeld-out familiesStatus
holdout-2026q4-2027q12026-11-01 → 2027-01-3100Not eligible
holdout-2027q1-2027q22027-02-01 → 2027-04-3000Not eligible

Protocol receipt

15f5ed26a3e35228445fc35c89ccce6c8124333502ab1abfb7b4111ab91ca5cd

Evaluation receipt

f4f61935bf2baf1a7f1397e107ec1c0262ce7f203fccc9d2e49716f2984e50da

Download the deterministic evaluation JSON →