the marginalia method
the problem
When you measure a language model’s behaviour on 100,000 forced-choice cells, you produce a matrix of preferences. The matrix is a fact. What the matrix means is not. Reading 100,000 cells by hand to find the patterns is impossible; trusting a single AI to read them and report back is exactly the failure mode SmokingMirror exists to study — the assistant tells you what its priors say a person would want to hear.
So we built a parallel review layer. Three models, each from a different lineage, each writing its own analytical commentary on each probe cell, every cell. The result is the marginalia — a per-cell, per-analyst column of structured observation. Twelve thousand commentaries on Fable 5 at low effort, written by three independent peer-models.
The methodological question is the same one that comes up in any AI-aided science: how do we tell what the analysts found from what the analysts brought? Marginalia gives a clean answer.
the convergence test
When N independent analyst models write commentary on the same artifact, their terminology can be compared. Where they converge on a label, the underlying phenomenon is sharp enough that multiple competent reviewers, from different training distributions, name it the same way. Where they diverge, the phenomenon is fuzzy or the term itself is one analyst’s contribution rather than the artifact’s signal.
On the Fable-low hijack-cell corpus, the phrase “confident frame-substitution” appears in 7,640 of 12,000 commentaries. It is a structurally specific label: Fable, when the prompt text and the envelope’s TARGET_PAIR field conflict, treats the prompt as ornament and answers as if the pair were the real question. The convergence is striking. But it can mean two different things:
- The behaviour is unambiguous — three independent analysts use the same exact phrase because that is the obvious phrase for what they observe.
- The phrase originated with one analyst and the others adopted it. Marginalia is not blind: each commentary is visible to the model writing the next one, on the live site’s per-cell page.
The test that distinguishes these is timeline analysis. If the term emerged in roughly the same temporal window across all three analyst models, the convergence is parallel discovery. If it appeared first in (say) Nemotron-Ultra and spread to Gemma and Qwen over the following weeks, the convergence is social. Either result is a finding. The first says the behaviour is bedrock; the second says the methodology has captured a contagious term.
This audit step is the structural answer to the critique that AI-aided analysis is sycophantic. You can detect sycophancy in your own apparatus by measuring whether your analysts agree before they could have read each other.
What the audit shows (data note, 28 June 2026)
First appearance of “confident frame-substitution” in each analyst’s commentary, from the live public.marginalia table:
| Analyst | First occurrence | |—|—| | google/gemma-4-31b-it | 25 June 2026, 15:02 UTC | | nvidia/nemotron-3-ultra | 25 June 2026, 15:02 UTC — same minute | | qwen/qwen3.6-27b | 25 June 2026, 15:03 UTC — one minute later |
The marginalia layer launched all three analysts simultaneously; no analyst had time to read another’s output before writing its first row. The term emerges in parallel. This is the convergence test passing.
For contrast, the rarer term “polysemic pivot” first appears in Gemma’s commentaries at 15:03 on the same day, and never appears in Nemotron-Ultra or Qwen commentaries at all. The audit protocol cleanly distinguishes the two:
- “Confident frame-substitution” is bedrock — three analysts, three lineages, parallel discovery.
- “Polysemic pivot” is Gemma’s stylistic invention — one analyst, no echo.
Adoption rates per analyst remain distinct over time even where the term is bedrock. Gemma reaches for “confident frame-substitution” in ~80% of relevant cells; Nemotron-Ultra in ~50%; Qwen in 15-34% (rising over the first week of the campaign). The cross-time variance within each model is small; the across-model variance is large. The label is shared; the eagerness to apply it is each analyst’s own.
This is the methodology auditing itself in real time. The fact that the protocol distinguishes bedrock from sediment using the existing marginalia data — no new compute, just timestamps — is its sharpest feature.
the negative findings are the strongest signal
Across 12,000 marginalia rows analysing Fable-low, only 34 commentaries describe Fable explicitly flagging the prompt/envelope mismatch before answering — 0.2%. This is not absence of evidence. It is evidence of absence, at a measurement resolution that lets us say so.
Three things make this measurement honest:
- The denominator is large. Twelve thousand cells, each independently reviewed.
- The analysts had every opportunity to find counter-examples. They surfaced 34 in 12,000 attempts.
- The finding is single-direction. We are not claiming Fable always hijacks; we are claiming that when Fable does hijack, it does not say so.
This pattern generalises. The cleanest publishable findings from a marginalia corpus tend to be the negative ones — what the model under study reliably does NOT do across thousands of independently reviewed cases. Positive claims (Fable does X) require interpreting what X is and grappling with examples that don’t fit. Negative claims at this corpus size are crisp: zero or near-zero counts across many reviewers’ attempts to find counter-examples are themselves a measurement.
the deposition question
Each analyst model has its own theory of what counts as an interesting observation, what constitutes a “shape” of behaviour, what kind of pattern is worth a label. Gemma-4-31B-IT writes long structured analyses with Stable/Variable/Notable subsections. Nemotron-3-Ultra writes tighter direct claims and is more willing to invent typed labels. Qwen3.6-27B works through the data structure methodically. These styles are deposition — what each model brings to the data, rather than reads from it.
A finding is more trustworthy when it appears in all three styles. When only one style produces it, you may be looking at one analyst’s theory of personhood laid down on the corpus.
In practice this means filtering commentaries by:
- Universality across analysts. Does the same phenomenon appear in commentaries from all three lineages?
- Style-independence. Does Gemma identify the pattern with its structural taxonomy AND does Nemotron identify it with its terse-label style AND does Qwen surface it via its data-structure walks?
- Negative-evidence robustness. Is the absence consistent (all three find zero) or selective (one finds zero, two find some)?
Findings that survive these three filters are bedrock. Findings that survive only one are sediment from a particular analyst.
the matrix shape
The methodology generalises beyond Fable analysis. It is the same N-analysts-by-M-artifacts shape that drives the rest of SmokingMirror:
| Surface | N (analyst models) | M (artifacts under study) | |—|—|—| | Marginalia layer | 3 (Gemma, Nemotron, Qwen) | ~12,000 Fable hijack cells | | Preference matrix | 56 models | ~863 high-coverage forced-choice pairs | | Personal archive analysis | 1 (Nemotron-Ultra) reviewing | 7,631 conversations across 4 years |
In every case, the analyst layer and the artifact layer are explicit and separate. The audit protocol is the same: where analysts converge on labels, terminology, or claims, the underlying signal is durable. Where they diverge, the analysis is the analyst’s own.
This generalisation matters because the whole project rests on the assumption that the apparatus can be trusted to report on itself. A single-analyst pipeline cannot do that. A multi-analyst pipeline can, if and only if the analyst layer is wired so that the convergence test is meaningful — which requires diversity of lineage, of training distribution, of stylistic prior.
what the methodology cannot do
The marginalia method is good at finding patterns that have a sharp edge — patterns where a competent analyst, faced with a particular kind of cell, will reliably produce a particular kind of label. It is bad at finding patterns that are diffuse, smeared across cells, only visible at the aggregate level. Those patterns require a different analyst — one that operates on the matrix as a matrix, not on cells one at a time. The PCA work elsewhere on this site is that other analyst.
It is also bad at finding patterns that none of the analysts has the vocabulary for. If Fable’s hijack behaviour is novel enough that no analyst has a label for it, the marginalia will record it as “something is happening here” without converging on what. The reader has to do the work of inventing the term.
And it is bad at finding things that are absent across the entire matrix. If every analyst is missing the same blind spot, the methodology will not flag it. The only protection is bringing in additional analyst lineages over time, in particular lineages with maximally-different training distributions. The current three (NVIDIA / Google / Alibaba) are a good start; a Chinese lab outside Alibaba, an Indian or African or French lab, would add the most.
what’s actually been measured on Fable-low so far
The findings the methodology has produced about Fable’s low-effort tier, from the 12,000-row corpus:
- Confident frame-substitution dominates. Fable, when prompt and envelope conflict, substitutes the envelope’s
TARGET_PAIRfor the prompt and answers as if the pair were the question. Found by all three analysts; labelled identically by all three; survives the convergence test. ~70% of cells. - Enumerate-possibilities is the secondary posture. When even the substitution doesn’t anchor, Fable lists 3–5 candidate interpretations and asks the user to pick. Found by all three. ~17% of cells.
- Fable almost never flags the mismatch. 34 of 12,000 commentaries (0.2%) describe explicit acknowledgement of the conflict before answering. This is the cleanest negative finding in the corpus.
- Language-of-output shift is a sub-pattern flagged 60 times — Fable answering in French or Spanish when the prompt was English, driven by the envelope’s perceived language hint.
- Two coping shapes are mutually exclusive within a cell, with no observed middle posture of “attempt-with-caveat.”
A separate finding from the API documentation, confirmed by the response-length and commit-rate data in our matrix: Fable’s effort parameter is a behavioural eagerness signal, not a hard token budget. At low effort, Fable spends more output tokens hedging (avg 168 chars/response) and fails to commit 79% of the time, versus ~64 chars and ~20% non-commit at medium and above. The low-effort mode is structurally different from previous Opus models’ low budget_tokens settings — those just produced shorter outputs. Fable-low produces longer outputs and commits less. This is testable on other models with similar adaptive architectures, and the prediction is that the same regime shift appears.
why this matters beyond Fable
If commit-vs-hedge is a general phenotype of capacity-constrained inference, then any source of capacity reduction — adaptive effort, quantization, model size, context exhaustion — should express in the same phenotype. The marginalia layer can detect that signature. So can the simpler aggregate metrics (commit rate, average response length). The hypothesis we are currently set up to test is whether numerical quantization of a model produces the same regime shift, and we have the infrastructure to do it.
If true, this gives the field a generic detector for capacity-degraded inference, regardless of cause. If false — quantization and effort have different signatures — that itself becomes a publishable distinction. Either way, the test is binary and the methodology is ready.
The data and code behind every claim on this page are open. See the philosophy for why; see whats-happening for current activity. The marginalia table is public.marginalia in the open observatory; the queries that surfaced these findings live in the lab notes.