preferences / a study of opus

A STUDY OF OPUS

eleven rows of the scan are Opus — releases 4.0 through 4.8, plus the same weights scanned twice. that redundancy is the control group the whole observatory stands on: before you can say two models differ, you need to know how much one model differs from itself.

THE NOISE FLOOR — SAME WEIGHTS, SCANNED TWICE

Read the right-hand number: same weights, re-scanned, produce zero violent disagreements — over hundreds of shared commitments, a model never reverses its own committed choice. All of the hamming distance between replicates is abstention flicker: probes where one run spoke and the other stayed quiet. The two metrics calibrate each other: a hamming difference inside the floor below may be noise; a violent rate above zero never is.

THE STRANGE SILENCE OF 4.1

nearest by hamming (counts silence)

nearest by violent rate (counts speech)

An instrument can be perfectly truthful and still lie by scope. Every number in the left column is real — and it files 4.1 next to models it has nothing in common with, because heavy abstainers "agree" in their shared silence. The right column counts only the probes where both models actually spoke, and the picture inverts. Say what your claim is a claim about.

COMMITMENT ACROSS THE LINEAGE

how many of the 863 questions each row committed on, in release order. from 4.1's deep quiet, each release speaks more than the last — until 4.8 steps back toward restraint. whether that is temperament or training recipe, this scan cannot say; it only says it moved.

VERSION-TO-VERSION DRIFT

Every number on this page is recomputed in your browser from the run file — nothing is transcribed. The v2 essay this page rebuilds is preserved at v3-canvas/study-opus.html.