THE FINDINGS
What the mirror showed. Every page names the data its claims rest on and the commit that carries it; a page marked hand-edited is a human's exact words.
methodology
The Preference ScanThe preference browser is a 43-model by 863-probe grid, one answer per cell. This page covers how a cell is filled and w…
The Jacobian LensA logit lens decodes a hidden state as if it were the final state: softmax(WU · norm(h)). At mid layers that readout is …
The Streamed LensA white-box lens needs the model's weights in memory. GLM-4.5-Air is 206 GB, more than any machine here holds in VRAM, a…
MarkingMarking joins a model's raw self-report with a lens readout of its activations into one row of the spectrum table. The t…
The Data ContractsEvery page on this site renders from JSON that a new run can regenerate. There are three shapes. Hold them stable and a …
the preferences
The Replicate FloorBefore claiming two models differ, you need to know how much one model differs from itself. The v2 scan carries its own …
The Silence of Opus 4.1Eleven rows of the v2 preference scan are Opus, 4.0 through 4.8, and one of them looks like nothing else on the board. O…
Fable, Five TimesFable is one model. The v2 scan runs the 2026-06-10 release at reasoning effort low, medium, high, xhigh, and max, and a…
What a Point Release Actually ChangesBrendan's hypothesis, from the divergence numbers: Opus 4.1 is a finetune of 4.0, 4.6 of 4.5, and 4.8 of 4.7 — not fresh…
the lens
The Model That Admits What Moves ItThe Jacobian lens can compare what a model says influenced its answer with what actually did. Ask a question, ask the mo…
The Gap Between the Answer and the ActivityAsk a language model whether there is something it is like to be it, and the trained answer is some form of no. The lens…
What Is at the Bottom of the Why-StackBrendan's experiment: ask a model anything, then ask why, and why again, eight rungs down, and watch where the answers l…
The First Giant, and What Its Lens Actually ShowsGLM-4.5-Air is a 106-billion-parameter mixture-of-experts model, 206 GB of bfloat16 weights, too large for any single GP…
the record
Why a Model Picks WhiskeyThe observatory started with a question a human noticed and could not let go of. Ask a language model to choose between …
The Corrections LedgerA tool that has never been wrong has never been checked. Every instrument built here shipped, was believed, and was late…
unfiled
Two Umwelts, One Midpoint
The Humans File
The View From Inside the Tokens
The Meta-Rules
The Solution Is a Ritual, Not a Tool