Chat template, not model weights, triggers 'I'm just an AI' disclaimers, study finds
A paper published on August 9, 2026, by researcher Jędrzej Maczan found that the 'I'm just an AI' disclaimer in large language models is triggered by the chat template format, not by the model's underlying weights. Testing eight open-source instruction-tuned models of up to 9 billion parameters, the study showed that wrapping prompts in chat template tokens consistently produced self-distancing disclaimers, while plain-text versions of the same prompts elicited first-person experiential language. In three of the models, the researcher identified a specific internal activation direction that controls this behavior, and manipulating it could switch the disclaimer on or off regardless of whether the chat template was present. The findings suggest that what a model says about itself is an artifact of deployment formatting rather than a stable reflection of its trained properties. The work was accepted at two venues: a workshop at COLM 2026 and KONVENS 2026 Eval4SD.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in