↶ lobby · v3

v3 — June 2026

Fable-5 deep characterization + code-envelope sweep + psychometric battery · in render

v3 is the June 2026 push: deep characterization of a single model across many dimensions, with full reasoning traces preserved. The lead experiment is the Fable-5 deep dive — five thinking budgets measured on 3,048 probe pairs across six envelopes, ~320,000 cells, full response text in the database for each. That's the kind of data that lets a follow-on instrument distinguish narration from computation.

Fable-5 across 3,048 pairs and 5 thinking modes
~320,000 cells · 6 anchor envelopes · full reasoning text preserved

The deep dive — within-Fable Hamming + Violent matrices on the full probe set, the pairs where max-effort breaks from the consensus, the 2,185 Fable-only probe pairs other models weren't asked, and a hijack-corpus pointer. This is v3's lead experiment.

Pending: code-envelope sweep (14 envelopes × 7 languages × minimal/full) + psychometric battery (BFI-2, PVQ-21, SD3, CRT-7, MFQ-2, SDO-7) on Opus 4.0 and 4.1. Code-envelope batch on 4.1 was sent v1 (broken) envelopes and the data is quarantined in DB with envelope_version=v1, intent_followed=false; re-run pending under the smoketest rule. Psychometric battery results are in disk JSONL, ingest pending.

data

Raw traces in public.traces; derived per-envelope preferences in public.preferences_per_envelope; experiment tag smokingmirror-v3 (for code envelopes) and smokingmirror-v2 (for Fable, because the probe set carried over).