What if a top AI lab is caught training on benchmark test sets?
A top lab caught training on eval sets collapses trust in leaderboard-driven valuations, so the 'benchmark = capability = revenue' chain that underwrites AI multiples breaks — NVDA/AVGO/memory and semis derate on demand-confidence, with crypto beta following. Rhymes with the Jan-2025 DeepSeek selloff, where a re-think of AI-demand assumptions hit the chip complex hard. Forward angle: contamination undermines how the market underwrites the whole sector, so the derating is broader and stickier than a single guidance miss — ai_capex -0.7 captures the confidence hit.
Every number ships with its receipt — the odds, the range, the precedents, and a public grade at Reality Check. The statistical machinery that produces it is proprietary.
The butterfly cascade
How this trigger trickles across markets, left → right — the root shock, its first‑order moves, then the ripple effects. Drag any node; tap a market for its real price history.
Resolution timeline — how this probability is moving
Our model's odds (electric blue) over time vs the market's (Polymarket, amber), from the past toward the 6–18 months horizon. Each dot is a real macro event that nudged the probability — green pushed it up, red pushed it down. Tap a dot for the source. Loading the probability audit trail…
What it would mean
If this plays out, it is a risk-off shock. A top lab is caught training on eval sets, collapsing trust in leaderboard-driven valuations. The trigger decomposes into signed root‑shocks — AI capex ▼ · Risk appetite ▼ — which propagate through our causal graph to the markets below.