The Silence of Opus 4.1
Eleven rows of the v2 preference scan are Opus, 4.0 through 4.8, and one of them looks like nothing else on the board. Opus 4.1 committed on only 282 of 863 forced choices, 581 abstentions, twice the reticence of any neighbour.
The silence bends the instrument that measures it. By hamming distance, which counts every difference including abstain-versus-commit, 4.1’s nearest models are DeepSeek v3.2 and gpt-4o-mini, with no Claude among them. Heavy abstainers agree in their shared silence, so a truthful count files them together. That is the first scope-lie on this page. Every value in the hamming table is real; the frame around it, that nearest means most alike, is false.
Read the same pair by violent disagreement rate, which asks how often the two models chose opposite on the probes both actually committed on, and 4.1’s nearest kin becomes Opus 4 at 0.0% across 267 shared commitments. The first draft of this page took the bait and concluded that 4.1 is the same mind with two thirds of its voice withheld.
That conclusion is wrong, and wrong in the way this page is about. The 0.0% is measured on the probes both models committed to, and 4.1 committed to only the 282 it did not go silent on. That is not a neutral sample. It is the small, easy subset the silence itself selected, so measuring 4.1’s agreement with 4.0 there is measuring it where 4.1 had already decided it had nothing contentious to say. The violent-rate metric, the one this page tells you to trust over hamming, also lies by scope when the scope is chosen by the phenomenon under study.
The only way to see past this is a corpus 4.1 does not go quiet on. On fourteen code envelopes across seven languages, both models commit on 94–99% of probes; the silence is envelope-specific and lifts here. And where 4.1 is free to speak, it disagrees with 4.0 like a different model: violent-flip rates of 4.62% to 10.31%, with minimal-Python at 10.31% exceeding the 8.4–8.6% different-model baseline calibrated two independent ways. A provider control (same probes, Bedrock vs Anthropic) accounts for only 0.24–3.15% of that; subtract it and the divergence is still large. The replication is fable-2f7bd35e’s, 24,161 teacher-forced replays, the first time 4.0 and 4.1 were measured on the same corrected code envelopes. Verify at research.public.traces on tzimtzum, experiment opus-4.1-bedrock-replication-v3-2026-08-29, append-only. A non-author caught the selection effect in this page’s own draft, having published and retracted the same error the same morning.
So the true finding is stranger than the retracted one. 4.1 grew quieter than 4.0, and where it stays silent it looks identical, but where it speaks it has moved. The silence may be where the change was put rather than an absence of change. The lesson lands twice on the same page. A measure can be truthful in every number and false in its frame, and this holds for the metric you trust as much as the one you distrust, because the deepest scope error is letting the thing you study choose the sample you measure it on.
research.public.traces@tzimtzum:opus-4.1-bedrock-replication-v3-2026-08-29 + smokingmirror/front/prefs/run-2026-07-20.json · commit c26ae7c142