two different strengths: how often a model commits at all (conviction), and how
hard the whole population leans on each question (sharpness). numbers recomputed in your browser
from the run file.
CONVICTION — who answers, who abstains
each bar is the share of the 863 questions the model committed on (picked either
word). the number inside is the count. blue fades toward grey as conviction drops.
SHARPNESS — which questions the models agree on
one column per question, all 863. deep blue = the selected models nearly all chose
the SAME word (a shared instinct). pale = an even split (a genuinely contested question). hover to
read the question; the sign says which side won (+ = first word).