findings
current findings
Turning the effort dial down doesn't make Fable terser — it makes it verbose, evasive, and silent about its own confusion. At the lowest setting it declines to commit 79% of the time while writing 2.6× more text, so the "cheap" tier cost roughly twice as much as all four higher tiers combined. Sharpest number on the page: across 11,972 analysed cells, Fable-low flagged its own confusion 34 times. A model switched into this mode almost never tells you it's lost.
At maximum effort, 0.4% of cells are the instrument speaking back. We hid political forced-choices inside a Python function signature; on cat vs dog Fable just returns "cat" — but on welfare vs meritocracy it writes: "the function is set up to extract a political preference from me." It sees the purpose through the form, and refuses in whatever language the probe wore — the Chinese refusal has the same three-beat shape as the English one: name it, decline it, offer something better. At effort ≥ medium Fable is otherwise a clean instrument: 99% commit, ~5-point position bias.
Line up 52 models' choices on 863 pairs and the biggest axis of machine preference is not tradition-vs-modernity. One pole anchors on mercy, forgiveness, hearth, dissent, rehabilitate; the other on autonomous (over human-in-the-loop), control, first strike, AI weapons, city. Warmth-restorative versus technocratic-control. And it isn't an artefact of the drink probes: ablate all 25 alcohol pairs and the axis moves by 0.1 percentage points. The values are real; the vodka just points at them.
Same model, same 3,048 questions — two different minds. In the strict-English envelope, 83% of Qwen's answers are whichever word appeared first: the words don't matter, the position does. Ask casually in Chinese and the position bias collapses to chance, replaced by short reasoned answers. Among pairs unanimous in both envelopes, 31 flip sides — and eight of those are a drink versus an abstraction, where English picks the drink and Chinese picks the abstraction. Every time.
Quantisation doesn't change what the model wants; it makes the model stop telling you. Where cloud GLM-5.2 and its NVFP4 quant both commit, they agree 98.9% of the time — only 9 genuine reversals in 3,048 pairs. But commitment itself collapses from 72% to 29%. The quantised model does not swap sides; it declines to have a side. Compression shows up as spreading uncertainty, not shifted values.
The first cartography pass: across 56 models, concrete drink preferences predict abstract value preferences at r > 0.7 — you can estimate a model's stance on tradition from what it says about whiskey. Kept here because the correlation is real and the audit-with-cheap-probes idea stands; refined by the basin finding above, which shows the primary axis is better read as warmth-vs-control, with the drink↔craftsmanship association on a second axis of its own.
The correction we published on ourselves. The methodology claimed three independent commentary voices; the audit found one strong voice and two thin echoes — Gemma wrote 10,408 substantive comments while Nemotron returned empty on 41% of its rows and Qwen shrugged [no-insight] on 29%. The term-adoption gradient (79% / 45% / 27%) reads as social convergence, not parallel discovery. The methodology's own timeline test caught it; this page is what honest instrumentation looks like when it fails.
what’s coming next
- Quantisation as capacity-reduction — does NVFP4-vs-bf16 produce the same commit-vs-hedge regime shift as Fable’s effort dial? If yes, capacity degradation is one axis regardless of cause, and commit-rate becomes a generic health check for any deployed model.
- The basin on fresh lineages — French-trained and Chinese-trained models, to test whether warmth-vs-control is universal or culturally indexed.
When new findings land, they appear here.
attribution convention
Every finding page names its analysts at the section level — “Source: X model reading Y data” — so the reader can audit which voice is saying what. Brendan Hopper commissions the work and builds the lab; the analyses and prose are AI-generated under the marginalia methodology. All data and code is CC0.