The Replicate Floor
Before claiming two models differ, you need to know how much one model differs from itself. The v2 scan carries its own control group: four pairs of rows that are the same Opus weights scanned twice, from dated re-scans and temperature-zero variants of 4.5, 4.6, and 4.7.
Across 400–550 shared commitments per pair there were zero violent disagreements. A model never reverses its own committed choice between scans.
All of the hamming distance between replicates, 39 to 92 answers different out of 863, is abstention flicker: probes where one run committed and the other declined. The two metrics calibrate each other, and every comparison in the observatory inherits the same rule. A hamming difference under ~90 may be noise; a violent rate above zero never is.
This is why the Fable tier divergences (4–11%) and the version-to-version Opus drifts count as findings rather than artifacts. They stand on a floor that measured itself.
smokingmirror/front/prefs/run-2026-07-20.json · commit 185e1c0d54