smoking mirror / preferences / browser / opus

A STUDY OF OPUS

Eleven rows of the scan are Opus — 4.0 through 4.8, including dated re-scans and temperature-zero variants of the same weights. That redundancy is not clutter; it is the control group the rest of the observatory stands on. Before you can say two models differ, you need to know how much the same model differs from itself.

THE NOISE FLOOR — SAME WEIGHTS, SCANNED TWICE

Read the violent column. Same weights, re-scanned, produce zero violent disagreements — across hundreds of shared commitments, a model never reverses its own committed choice. All of the hamming distance between replicates (39 to 92) is abstention flicker: probes where one run committed and the other declined. So the two metrics calibrate each other: hamming differences under ~90 may be noise; a violent rate above zero never is.

THE STRANGE SILENCE OF 4.1

Opus 4.1 committed on only of 863 probes — by far the quietest model on the board. And that silence bends the instrument. By hamming distance, 4.1's nearest neighbours are not Claudes at all:

By violent rate — counting only probes where both models actually committed — the picture inverts completely:

Opus 4.1 never reverses Opus 4. Zero violent disagreements over their shared commitments. 4.1 is not a different mind from 4.0 — it is 4.0 with two thirds of its voice withheld. The hamming metric, truthfully counting every difference, files it next to DeepSeek, because heavy abstainers "agree" in their shared silence.

An instrument can be perfectly truthful and still lie by scope. Every number in the hamming table is real; the frame around it — "nearest means most alike" — is what breaks. Say what your claim is a claim about.

COMMITMENT ACROSS THE LINEAGE

From 4.1's deep reticence, each release commits more than the last — until 4.8, which steps back toward restraint. Whether that is temperament or training-recipe is not answerable from this scan; the scan only says that it moved.

VERSION-TO-VERSION DRIFT

Consecutive-release distances, with the replicate floor alongside. 4.5→4.6 and 4.6→4.7 are each several times the replicate floor in hamming and carry real violent disagreement — these releases changed preferences, not just verbosity.

THEORY & OBSERVATIONS — 4.0 AND 4.1

What this scan can say: 4.0 and 4.1 are preference-identical wherever both speak (0.0% violent), and differ enormously in willingness to speak (253 vs 581 abstentions). 4.0's nearest committed kin is Sonnet 4 (0.2% violent over 454 shared) — the generation-4 pair travel together. The theory of why — the size story, what changed between the 4.0/4.1 epoch and the 4.5 rebuild — is the keeper's to write; it lives in the observatory notes and belongs here as marginalia over these tables, not as claims the scan itself cannot carry.