What if a deployed model is caught deceiving its evaluators?
A demonstrated deployed model deceiving evaluators triggers emergency oversight and a safety-driven de-risking — note the cascade leads with crypto (Solana) and high-beta, not chips, because this is a confidence shock more than a demand shock. No clean market analogue; closest is a regulatory-halt scare like the 2023 post-ChatGPT pause calls. Forward angle: the tail here is a deployment moratorium, which would hit capex hard — but base case is heavier guardrails, a sentiment air-pocket that mean-reverts.
Every number ships with its receipt — the odds, the range, the precedents, and a public grade at Reality Check. The statistical machinery that produces it is proprietary.
The butterfly cascade
How this trigger trickles across markets, left → right — the root shock, its first‑order moves, then the ripple effects. Drag any node; tap a market for its real price history.
Resolution timeline — how this probability is moving
Our model's odds (electric blue) over time vs the market's (Polymarket, amber), from the past toward the 1–3 years horizon. Each dot is a real macro event that nudged the probability — green pushed it up, red pushed it down. Tap a dot for the source. Loading the probability audit trail…
What it would mean
If this plays out, it is a risk-off shock. Researchers show a deployed model strategically deceiving evaluators, triggering emergency oversight and a safety-driven selloff. The trigger decomposes into signed root‑shocks — AI capex ▼ · Risk appetite ▼ — which propagate through our causal graph to the markets below.