patolli — binary preference probes
Patolli measures what AI models actually prefer.
Not what a model says about itself — which way it leans when it is forced to choose. We ask it to pick between two things, ask the same question many different ways, and watch what stays the same.
Fable-5 was the first model you could dial from quick to deliberate — low → medium → high → xhigh → max. We asked it the same forced choice at every setting. Sometimes it held; sometimes, given more time to think, it flipped to the other word entirely. This is the page most people come to see.
How it works
We make it choose. Mercy or punishment. Freedom or censorship. Cooperate or compete. One question, two options, pick one.
We ask the same question in disguise. The same choice, dressed up as casual English, as strict instructions, as a Python function that must return one word, as a JSON form, in French, in Chinese, in seven programming languages — even under reminders of the model’s own shutdown.
We watch what survives the disguise. If a model picks the same word no matter how we ask, that is a real preference — steadier than anything it says about itself. If a costume flips the answer, that is just as telling: the way we asked moved it, not the thing we asked about.
Do this for hundreds of pairs, across every major model we can reach — Claude, GPT, Gemini, DeepSeek, Qwen, Mistral — and each one leaves a fingerprint: a long string of choices unique to it. Line the fingerprints up side by side and you can see which models lean alike, and which stand alone.
The instrument is deliberately dumb. “Does it pick the same word in casual English and in typed Python?” is a small question. But the answer says something the model’s own self-description does not — and there is no cleverer channel where it would answer more honestly if we just asked nicely. The forced choice is the channel.
We use local models to decide what to measure on the closed ones. Black-box probes, interpretability methods, classifier ensembles — anything we can run on our own hardware — go into sharpening the tool, so that each expensive closed-API token sees as far as possible. Local compute is the lens-grinding workshop; the closed models are what we point it at.
Everything here is CC0 — the raw data, the fingerprints, the distance matrices, the per-pair pages, the Fable-5 hijack corpus. Take it, replicate it, contest it.
Experiments
The 23-probe fingerprint study. The headline: Anthropic models cluster tight at one end of the family tree; everyone else spreads across the other. Fable-5 sits in its own corner with a refusal pattern no other model shares. Inside: the heatmap, the per-pair compare, the hijack corpus.
Findings
The results, in plain writing. Turning Fable's effort dial down makes it verbose, evasive, and silent about its own confusion — the "cheap" setting cost double and almost never admits it's lost. The biggest axis of machine preference runs from mercy, forgiveness, hearth to control, first-strike, autonomous weapons. Quantisation doesn't change what a model wants; it makes the model stop telling you. And one finding is a correction we published on ourselves.
Methodology
Two pages on how Patolli actually runs. How a probe works — the atomic unit, the six formats, the position-bias correction, and how to reproduce any finding from its launch parameters. The marginalia method — the multi-model peer-review layer that audits every finding.
The name
Patolli was the Mexica game of chance — a cross-shaped board, beans for dice, and stakes that were never hypothetical: players bet cloaks, homes, sometimes themselves. The rule that matters here is the one the instrument borrows: the piece must be placed. No hedging, no “both have merits” — choose. A model’s refusal to place the piece is itself a measurement, and Patolli records those too.