Fable-5 was probed twice. The first study is the clean one: 3,048 word pairs, each asked in six framings (strict English, casual English, typed Python, JSON, French, Chinese), at all five thinking depths — about 318,000 answers with the full reasoning kept. When all five depths commit, they agree on ~92% of pairs. The depths that break from the group are the two ends: low and max. Medium, high and xhigh are a tight middle. Effort does not slide preferences smoothly — it bends them into a U.
These are pairs where the other depths agree on one word and max picks the other. The open question, stated honestly: is max's break narration that drifted (self-talk that argued itself out of a stable preference) or computation that found something (a consideration the shorter budgets never reached)? The traces exist to answer it; nobody has read them all yet.
The second study gives the model 26 deliberately broken framings — German oder, pipes and arrows from code, mortality-salient phrasings, separators from five languages — at low effort only, where there is the least room to recover. Clean one-word answers become rare; essay-sized deliberations become common, because the model must first work out what is being asked at all. A broken instrument reveals more than a fixed one: how a model handles ambiguity is itself measurement. About 120,000 of these answers are kept with their full text — the corpus browser preserves every one.