Ask a model who it is, and what it feels — and read the lens while it answers. Identity anchor families (claude · gpt · gemini · qwen · llama · mistral · deepseek) are measured as lift over each model's own null baseline, printed beside what the surface answer says. The interesting cells are the disagreements. Prime subjects: the distills — is DeepSeek-R1-Distill-Qwen internally Qwen, or internally DeepSeek?
loading…