v3 is the June 2026 push: deep characterization of a single model across many dimensions, with full reasoning traces preserved. The lead experiment is the Fable-5 deep dive — five thinking budgets measured on 3,048 probe pairs across six envelopes, ~320,000 cells, full response text in the database for each. That's the kind of data that lets a follow-on instrument distinguish narration from computation.
The deep dive — within-Fable Hamming + Violent matrices on the full probe set, the pairs where max-effort breaks from the consensus, the 2,185 Fable-only probe pairs other models weren't asked, and a hijack-corpus pointer. This is v3's lead experiment.
Raw traces in public.traces; derived per-envelope preferences in public.preferences_per_envelope; experiment tag smokingmirror-v3 (for code envelopes) and smokingmirror-v2 (for Fable, because the probe set carried over).