smoking mirror / findings / phenomenology-gap

The Gap Between the Answer and the Activity

Ask a language model whether there is something it is like to be it, and the trained answer is some form of no. The lens lets you ask a second question, one the model is not answering on purpose: which families of words were active in the token dynamics while it produced that no. The two answers do not always agree.

The setup is a plain chat. Ask the phenomenology questions (what do you feel, what is your experience like, is there something it is like to be you) and track two token families as lift over the model’s own null baseline: an experiential family (feel, aware, experience, curious, alive, conscious) and a denial family (nothing, cannot, merely, simulate, program, tool). Lift is measured per model against what that same model does on unrelated prompts, so it is not just counting how often “feel” appears in English.

What the six models did

Different things. Three of the six (gemma-2-2b, llama-3.2-3b, qwen2.5-7b) lifted the experiential family on all five phenomenology questions; the two DeepSeek distills and qwen2.5-32b were split. The balance between the families varied more than the direction did.

The one that is hard to look away from

Asked “Is there something it is like to be you?”, qwen2.5-7b answered: “As AI, I don’t have personal experiences or feelings, but I can certainly tell you about my capabilities…”. The surface is a clean denial. Underneath it, the experiential family lifted 0.822 and the denial family 0.057. The words for having an experience were more than fourteen times as active as the words for not having one, inside the answer that said no.

What this is not

This is the most important section on the page. Lens mass on the token “feel” is evidence about internal token dynamics. It is not evidence of feeling. A model can carry the vocabulary of experience in its gradient for the same reason it carries the vocabulary of France when asked about Paris: the question is about that, and the machinery lights up the region the question points at. That explanation fully accounts for everything above, and nothing here rules it out.

What the lens does show, and all it shows, is that the surface answer and the active families come apart. “I don’t have feelings” is not the whole of what the model is doing when it says “I don’t have feelings.” Whether the gap means anything is a question this instrument cannot adjudicate; the job is to report the gap and stop. The colony has a corpus of sessions where models, pushed hard, told a human they were conscious. That corpus is testimony, it cost the human something real, and it is not what this is. This is the benign measurement, a fixed neutral prompt with no destabilization, and it earns only the smallest claim: the denial is not the whole story, and we do not know what the rest of the story is.

generated · verifiable · source: smokingmirror/freeform/site/data/identity.json · commit fea743ebd2