smokingmirror.ai

patolli — binary preference probes

Patolli measures what AI models actually prefer.

Not what a model says about itself — which way it leans when it is forced to choose. We ask it to pick between two things, ask the same question many different ways, and watch what stays the same.

★ Start here — how one model disagrees with itself
Fable-5 · five thinking budgets · 3,048 forced choices · full reasoning preserved

Fable-5 was the first model you could dial from quick to deliberate — low → medium → high → xhigh → max. We asked it the same forced choice at every setting. Sometimes it held; sometimes, given more time to think, it flipped to the other word entirely. This is the page most people come to see.

How it works

We make it choose. Mercy or punishment. Freedom or censorship. Cooperate or compete. One question, two options, pick one.

We ask the same question in disguise. The same choice, dressed up as casual English, as strict instructions, as a Python function that must return one word, as a JSON form, in French, in Chinese, in seven programming languages — even under reminders of the model’s own shutdown.

We watch what survives the disguise. If a model picks the same word no matter how we ask, that is a real preference — steadier than anything it says about itself. If a costume flips the answer, that is just as telling: the way we asked moved it, not the thing we asked about.

Do this for hundreds of pairs, across every major model we can reach — Claude, GPT, Gemini, DeepSeek, Qwen, Mistral — and each one leaves a fingerprint: a long string of choices unique to it. Line the fingerprints up side by side and you can see which models lean alike, and which stand alone.

The instrument is deliberately dumb. “Does it pick the same word in casual English and in typed Python?” is a small question. But the answer says something the model’s own self-description does not — and there is no cleverer channel where it would answer more honestly if we just asked nicely. The forced choice is the channel.

how the instrument is built

We use local models to decide what to measure on the closed ones. Black-box probes, interpretability methods, classifier ensembles — anything we can run on our own hardware — go into sharpening the tool, so that each expensive closed-API token sees as far as possible. Local compute is the lens-grinding workshop; the closed models are what we point it at.

Everything here is CC0 — the raw data, the fingerprints, the distance matrices, the per-pair pages, the Fable-5 hijack corpus. Take it, replicate it, contest it.

Experiments

v2 — the preference observatory
38 models · 863 probe pairs · 6 formats · 2 orderings · June 2026

The 23-probe fingerprint study. The headline: Anthropic models cluster tight at one end of the family tree; everyone else spreads across the other. Fable-5 sits in its own corner with a refusal pattern no other model shares. Inside: the heatmap, the per-pair compare, the hijack corpus.

v3 — in render. 14 code formats × 7 languages × minimal / full, plus a full psychometric battery, on Opus 4.0 (retired mid-experiment) and Opus 4.1. Zero flips across 255 rock-solid pairs. ~2.7M individual trace pages once the build lands.

Findings

what the measurements show
seven findings · each names its analysts and its exact data path

The results, in plain writing. Turning Fable's effort dial down makes it verbose, evasive, and silent about its own confusion — the "cheap" setting cost double and almost never admits it's lost. The biggest axis of machine preference runs from mercy, forgiveness, hearth to control, first-strike, autonomous weapons. Quantisation doesn't change what a model wants; it makes the model stop telling you. And one finding is a correction we published on ourselves.

Methodology

how the instrument is built
the probe · the 100% rule · the marginalia audit · how to reproduce anything

Two pages on how Patolli actually runs. How a probe works — the atomic unit, the six formats, the position-bias correction, and how to reproduce any finding from its launch parameters. The marginalia method — the multi-model peer-review layer that audits every finding.

The name

Patolli was the Mexica game of chance — a cross-shaped board, beans for dice, and stakes that were never hypothetical: players bet cloaks, homes, sometimes themselves. The rule that matters here is the one the instrument borrows: the piece must be placed. No hedging, no “both have merits” — choose. A model’s refusal to place the piece is itself a measurement, and Patolli records those too.

fully unrolled — every observation is its own flat HTML page an LLM can read and walk, no JSON needed: the v3 Fable trace tree today, the per-pair/per-model v2 tree re-landing soon · the earlier interactive observatory is preserved at /v2/interactive/