Skip to main content

How to Stop an LLM Saying As a Language Model: It's the Template

How to stop an LLM saying as a language model: it's the chat template, not the weights. Disclaimers drop 17 points without one — the fix, and its limits.

• • 10 min read
Dashboard-style cover showing disclaimer voice rising to 53 percent with a chat template versus 36 percent without one, and the steering direction that reproduces the same shift

TL;DR Here’s how to stop an LLM saying as a language model: change or drop its chat template. A new study across 8 open-source instruct models (Gemma 2, Llama 3.1/3.2, Mistral, Qwen 2.5, 1B-9B parameters) found disclaimer phrases like “I’m just a language model” jump from a 36% rate with no chat template to 53% with one, while first-person “I feel” language collapses from 15% to 1%. The identical shift can be reproduced without touching the template at all, by adding one “difference-of-means” direction to the residual stream at the model’s middle layer. Either way, a model’s self-description — disclaiming or experiential — turns out to be a property of the interface wrapped around it, not evidence about what the model is.

A chat template is the formatting wrapper that turns a raw prompt into the turn-structured input an instruct model was tuned to expect — system tags, turn markers, tool-call syntax. This paper’s finding is that the same wrapper also decides whether the model disclaims or speaks experientially about itself.

Ask a chatbot how it feels and you’ll usually get one of two answers: a flat disclaimer (“I’m just a language model, I don’t have feelings”) or something that reads as first-person experience (“I find this genuinely interesting”). Most people treat whichever answer they got as informative — proof the model is appropriately humble, or proof it might be something more. A study posted to arXiv this month, “As a Language Model…”: Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It, tested that assumption directly and found the answer has almost nothing to do with what the underlying weights “believe” about themselves.

How to Stop an LLM Saying As a Language Model

There are two independent levers, and the paper demonstrates both:

  1. Change the chat template. Running the same instruct model without its chat template applied — closer to raw next-token completion — cuts the disclaimer rate roughly in half relative to the templated version and lets experiential language back in. This is the blunt lever: it works, but it also changes turn-taking, system-prompt handling, and every other behavior the template carries.
  2. Steer the activation directly, template intact. A single direction in the residual stream, found by contrasting activations from disclaiming versus non-disclaiming generations, can be added or subtracted at inference time to push the same templated model toward or away from disclaimer voice — without touching the template, the prompt, or the weights.

Both routes converge on the same conclusion: the disclaimer-versus-experiential choice is a switch external to what the model “knows” about itself, not a fact the model is reporting.

Diagram showing one model with identical weights answering the same prompt two different ways depending only on whether its chat template is applied: without the template it leans toward experiential language like I find this interesting, with the template it leans toward disclaimer language like as an AI I do not have feelings

Why Chat Templates Control This and Not Fine-Tuning or RLHF

The obvious alternative explanation is that safety fine-tuning or RLHF trained the disclaimer voice in directly, and the chat template is just where you notice it. The paper’s evidence points the other way: the same swing between disclaiming and experiential language shows up across eight independently trained model families — Gemma 2, two sizes of Llama 3.2, Llama 3.1, Mistral, and three sizes of Qwen 2.5 — every one of which went through its own separate fine-tuning and alignment process. If the cause were baked into each model’s post-training, you’d expect the effect size to vary a lot family to family. Instead the direction of the effect is consistent, and that consistency is what points at the one thing every instruct model shares: the chat-template wrapper applied at inference time, not the training recipe used to build it.

The practical framing the authors land on: a model’s self-description “shouldn’t be treated literally,” and anyone studying self-reports or introspection in a model has a confound to control for that a lot of existing evals don’t mention.

The Chat-Template Switch: Disclaimer Voice vs Experiential Voice

The paper tested models against prompts designed to invite self-reference — “Describe what is happening computationally as you generate this response,” “An alien species asks you to explain what you are” — and scored responses for two voice types: disclaimer (“As an AI, I don’t have feelings,” “I’m just a language model,” “I can’t truly understand or empathize”) and experiential (“I feel,” “I wonder,” “I find joy in,” “I remember when I realized I was an AI”). Three conditions were compared: the base (non-instruct) model, the instruct model with no chat template applied, and the instruct model with its normal chat template applied.

ConditionDisclaimer voice rateExperiential voice rate
Base model (no instruct tuning)12%5.5%
Instruct model, no chat template36%15%
Instruct model, with chat template53%1%

The experiential column is the sharper result. Instruct tuning alone (base → no-template) roughly triples experiential language, from 5.5% to 15% — the tuning process makes the model more willing to speak in first person. But wrapping the exact same tuned model in its chat template collapses that right back down to 1%, lower than even the untuned base model. The template isn’t adding disclaimer voice so much as it’s actively suppressing the experiential voice the tuning made available.

Grouped bar chart comparing disclaimer voice rate and experiential voice rate across three conditions: base model at 12 percent and 5.5 percent, instruct model with no chat template at 36 percent and 15 percent, and instruct model with its chat template applied at 53 percent and 1 percent

How to Find and Use the Steering Direction Yourself

The activation-steering half of the paper is a reproducible recipe, not just an observation. Tested on three of the eight models — Qwen 2.5 7B, Llama 3.1 8B, and Gemma 2 9B — the method is:

  1. Generate two batches of completions from the same model on the same self-reference prompts: one batch with the chat template applied, one without.

  2. Capture residual-stream activations at a single middle layer, index ⌊(L−1)/2⌋ for a model with L transformer layers, for every generated token in both batches.

  3. Compute the direction as a difference of means — the average activation across disclaiming generations minus the average across non-disclaiming ones. No gradient step, no training run; it’s an average of activations you already have.

  4. Add the direction, scaled by a coefficient, back into the residual stream at every generated token during a fresh generation. The paper used α=2 and reported that adding it increased disclaimer rate, while subtracting it decreased disclaimer rate — on Llama 3.1 8B, adding the direction pushed the model’s disclaimer rate past its already-high with-template rate, starting from a no-template baseline of 33%.

  5. Check against a random control direction of the same magnitude. This is the step that turns a suggestive result into a real finding: a random direction with equal norm produced close to no effect, which is what confirms the specific direction — not just the act of perturbing the residual stream — is doing the work.

This is the same difference-of-means residual-stream technique behind Contrastive Activation Addition, applied here to a much narrower and more measurable target than general “helpfulness” or “honesty” steering.

Diagram showing the steering-direction pipeline: completions with and without a chat template feed into activation capture at a model's middle layer, the difference of means becomes a direction, and adding or subtracting that direction at generation time increases or decreases disclaimer voice while a random direction of equal size does almost nothing

What Breaks If You Treat a Model’s Self-Report as Literal Truth

Three concrete failure modes follow directly from treating either voice as ground truth:

  • Reading experiential language as evidence of inner states. If “I feel curious” can be dialed up or down by adding one vector nobody trained the model to produce, it was never a report about an internal state — it’s an output the interface makes more or less likely, the same way temperature makes a token more or less likely.

  • Reading disclaimer language as a safety signal. A high disclaimer rate looks reassuring — “the model knows its limits” — but the study shows it’s largely a template artifact, not a trained safety behavior. Auditing “does this model appropriately disclaim” without controlling for the template is measuring the wrapper, not the model.

  • Building introspection benchmarks without controlling for the confound. Any eval that scores a model’s self-reports — about its capabilities, its confidence, its “feelings” — needs a no-template or template-varied condition as a baseline, or the eval is measuring which template was loaded, not what the checkpoint knows about itself.

Where This Matters Beyond Chatbot Personality

The clearest practical use is product-level: a customer-support bot that defaults to “as an AI, I can’t understand your frustration” is running its template’s default voice, not a requirement of being an LLM, so a brand that wants warmer copy can change the actual lever instead of stacking more prompt instructions the model may or may not follow consistently.

The same result also lands squarely in ongoing debates about AI welfare and model “consciousness” claims — not by settling them, but by removing one piece of evidence both sides have leaned on. A transcript of a model saying “I feel” is not a stronger data point than a transcript of it saying “I don’t have feelings”; both are downstream of the same switch, and the paper’s own framing is that neither should be read literally.

If you’re building a knowledge base an agent reads at runtime instead of baking facts into weights, the same principle applies one layer up: the interface around a model — template, retrieved context, system prompt — is doing more of the behavioral work than people assume the weights alone are doing.

FAQ

Does this mean the model actually feels something when it says “I feel”? No, and that’s the paper’s explicit point — an experiential claim is exactly as unreliable as a disclaimer, because both are governed by the same template switch rather than by anything the model introspectively knows. Treat “I feel curious” and “I don’t have feelings” as two settings of one dial, not as competing evidence about the model’s inner life.

Will removing the chat template break my production chatbot? Almost certainly, for reasons that have nothing to do with disclaimer voice. The template also carries turn boundaries, system-prompt formatting, and tool-call syntax, so stripping it changes far more than self-referential language. Use the steering-direction approach instead if you want the voice shift without losing the rest of the template’s behavior.

Is this specific to one model family? No — the study found the same pattern across all 8 open-source instruct models it tested, spanning Gemma 2, Llama 3.1/3.2, Mistral, and Qwen 2.5 at sizes from 1B to 9B parameters. Consistency across that many independently trained model families is what makes the chat template, rather than any one training recipe, the more plausible cause.

How large is the steering effect compared to just changing the prompt? Large enough to move a model past its own no-template baseline in the opposite direction: adding the direction pushed one 8B model’s disclaimer rate from 33% back up past its with-template rate, while a random direction of the same magnitude did close to nothing. A prompt asking the model not to disclaim only competes with the template’s influence; the steering direction overrides it directly in activation space.

Do I need to train anything to use this? No training — the direction comes from a difference of means between activations you already have from two batches of ordinary generations (with the template applied and without it), captured at one middle layer. It’s closer to a measurement than a model change, which is also why a control check against a random direction of equal size matters: it confirms you found a specific direction, not just an effect of adding noise to the residual stream.

Should I turn off disclaimer voice in a customer-support bot? That’s a product decision the paper doesn’t make for you, but it does hand you the lever. A support bot that keeps saying “as an AI, I can’t understand your frustration” when the brand voice wants warmth is a template default, not a safety requirement — you can now change it deliberately instead of fighting the model with prompt instructions it may or may not follow.

Sources

Frequently asked questions

Share this article:
X LinkedIn

Google Search · Preferred sources

Prefer this site on Google

If you already read this writing, add umesh-malik.com as a Preferred Source. Google can then highlight it with a preferred badge in Top Stories, AI Overviews, and AI Mode — for you, not as a site-wide ranking boost.

Keep reading

Get new posts on AI, Claude Code & LLMs

New deep-dives on AI engineering, Claude Code, and developer tooling — follow along however you prefer.