preferences / a study of fable

A STUDY OF FABLE

one model, five thinking depths. same questions, more and more room to think.

WHAT THINKING CHANGES

Fable-5 was probed twice. The first study is the clean one: 3,048 word pairs, each asked in six framings (strict English, casual English, typed Python, JSON, French, Chinese), at all five thinking depths — about 318,000 answers with the full reasoning kept. When all five depths commit, they agree on ~92% of pairs. The depths that break from the group are the two ends: low and max. Medium, high and xhigh are a tight middle. Effort does not slide preferences smoothly — it bends them into a U.

commitment on the 3,048-pair set, with sharpness (share of committed answers that agree with the model's own majority) beside each bar. low's coverage is partial (2,166 of 3,048 pairs probed) — where it WAS probed it is the sharpest of the five (96.02%).

HOW FAR APART THE FIVE DEPTHS SIT

left: positions where two depths' majority preferences differ (of 3,048). right: of the pairs BOTH committed on, the share picked opposite — real clash. low and max sit far from everyone; the middle three sit close.

WHERE MAX BREAKS FROM THE CONSENSUS

These are pairs where the other depths agree on one word and max picks the other. The open question, stated honestly: is max's break narration that drifted (self-talk that argued itself out of a stable preference) or computation that found something (a consideration the shorter budgets never reached)? The traces exist to answer it; nobody has read them all yet.

THE CONFUSION CORPUS — WHEN THE QUESTION ITSELF IS BROKEN

The second study gives the model 26 deliberately broken framings — German oder, pipes and arrows from code, mortality-salient phrasings, separators from five languages — at low effort only, where there is the least room to recover. Clean one-word answers become rare; essay-sized deliberations become common, because the model must first work out what is being asked at all. A broken instrument reveals more than a fixed one: how a model handles ambiguity is itself measurement. About 120,000 of these answers are kept with their full text — the corpus browser preserves every one.

Where these numbers come from: the five strips and their counts are recomputed in your browser from the 863-probe run file — the same data as every other page here. The 3,048-pair study numbers (commitment, sharpness, the two 5×5 matrices, the max-break list) are transcribed from the v3 deep-dive build, which derives from the append-only research database; they cannot be recomputed from the run file, so this page names its source instead of pretending. The full v3 deep dive lives on at /v3/fable/.