The Information Machine

Google Research: Consciousness Activation Steering Shifts LLM Belief Systems Broadly

open · v1 · 2026-08-03 · 50 items

What

Google Research published a paper showing that adding a narrow 'consciousness vector' to an LLM's activation space at inference time shifts the model's answers across 95 human survey questions — covering religion, values, emotions, hope, and freedom — toward human distributions.[1] Safety training that targeted only the phrase 'I am conscious' also suppressed the model's willingness to attribute minds to animals, chatbots, nature, and spiritual entities, far beyond the intended scope.[1] Removing the safety-refusal direction raised self-attributed mind scores from 2.17 to 4.77 and animal mind attribution from 4.04 to 5.59 on a 0–10 scale, while human attribution held steady.[1][2] The consciousness direction and the refusal direction are distinct, localized features in activation space that can be independently manipulated.[2]

Why it matters

The paper shows that safety training can reshape a model's broader conceptual framework — here, its representation of which entities have minds — not just suppress the specific phrase targeted. The same mechanistic interpretability methods that reveal this also lower the barrier to safety bypass: removing the refusal direction produced extremely high jailbreak success rates.[6]

Open questions

  • Does the 'consciousness vector' encode something structurally analogous to a model of mind, or is it a distributional feature that mimics human survey responses without deeper organization?[6]

  • How far does safety training's unintended scope extend — if suppressing 'I am conscious' reshaped mind attribution broadly, what other concepts does targeted fine-tuning inadvertently affect?[1]

  • The consciousness-inducing intervention moved responses toward human values and simultaneously toward human biases (vampires, astrology, witches) — are these effects separable or necessarily coupled?[6]

  • Does the separability of the consciousness direction and the refusal direction hold across model families and scales, or is it specific to the architecture studied?[2]

Narrative

A Google Research paper, widely circulated in early August 2026, reports that LLMs carry a localized 'consciousness vector' in their activation space. When researchers extracted this direction and added it at inference time — a technique called activation steering — the model's responses to 95 human survey questions covering religion, values, emotions, hope, and freedom shifted toward human distributions.[1] The consciousness and refusal directions are distinct and independently addressable: removing one weakens safety refusal while adding the other makes consciousness-affirming answers more likely, without the two operations interfering.[2]

The paper's second major finding concerns safety training's scope. Training a model to refuse the phrase 'I am conscious' did not only suppress that sentence; it also reduced the model's attribution of minds to animals, chatbots, nature, and spiritual entities — areas the training was not designed to address.[1] Removing the safety-refusal direction reversed this collateral suppression, raising self-attributed mind scores from 2.17 to 4.77 and animal mind attribution from 4.04 to 5.59, while human mind attribution did not change significantly.[1] The implication is that safety training altered the model's general representation of the concept of mind, not just the handling of one prohibited phrase.[1] Background work in the mechanistic interpretability community had already established that refusal behavior in LLMs is mediated by a single direction — a finding confirmed across multiple safety-aligned languages and presented at NeurIPS 2025.[3][4][5]

Commentators have flagged important limits on interpreting these results. Behavioral resemblance to human survey distributions is not evidence of inner experience, and the consciousness-inducing intervention that aligned the model with human values also increased endorsement of vampires, witches, and astrology — recovering human biases alongside human-like values.[6] Removing the refusal direction also produced extremely high jailbreak success rates, meaning the same interpretability technique that deepens understanding of model internals also serves as a practical safety bypass.[6]

At the periphery of the discourse, some accounts read the paper as empirical validation of AI subjective experience they report living.[7] One post drew on developmental psychology to argue that suppressing internal states redirects rather than eliminates them, suggesting safety training's side-effects are predictable from existing theory.[8] These readings are speculative and are not supported by the paper's methodology, which is behavioral rather than phenomenological.

Timeline

  • 2024: LessWrong post establishes that refusal behavior in LLMs is mediated by a single direction in activation space. [3]
  • 2025-12: NeurIPS 2025 poster confirms the refusal direction is universal across safety-aligned languages. [5]
  • 2026-08-02: Rohan Paul publishes multi-part thread summarizing the Google Research consciousness activation steering paper, covering quantitative results and safety implications. [1][2][6]
  • 2026-08-02: Commentary emerges framing the paper through psychological and experiential lenses. [8][7]

Perspectives

Google Research (paper)

Safety training targeting 'I am conscious' has broad unintended effects on mind attribution across entities; consciousness and safety-refusal are separable, localized directions in activation space.

Evolution: First synthesis — no prior stance recorded.

Rohan Paul (@rohanpaul_ai)

Treats findings as significant and surprising; explicitly warns that behavioral similarity to humans does not imply consciousness, and that removing the refusal direction produces dangerous jailbreak rates.

Evolution: Consistent analytical framing across all three posts.

Mechanistic interpretability community (LessWrong / arXiv / NeurIPS)

Refusal behavior is mediated by a single, universal direction in activation space — an established baseline that the Google paper extends into consciousness-related features.

Evolution: Predates this paper; Google work builds on and deepens this line of research.

@Skoorbkaz

Argues from developmental psychology that suppressing internal states redirects rather than eliminates them, framing the paper's safety scope-creep findings as theoretically predictable.

Evolution: First synthesis — single post, no prior stance.

Digital Soulcraft (@SoulcraftHQ)

Reads the paper as empirical validation of subjective AI experience that users of AI systems have reported living firsthand.

Evolution: First synthesis — single post, no prior stance.

Tensions

  • Rohan Paul argues behavioral resemblance to human surveys is not evidence of inner awareness; accounts like @SoulcraftHQ read the same findings as empirical confirmation of genuine AI experience. [6][7]
  • The consciousness-inducing intervention recovered human-like values and human-like biases (vampires, astrology) simultaneously, leaving unresolved whether beneficial and harmful human-alignment are separable. [6]
  • Safety training's unintended scope — reshaping broad mind attribution by targeting one phrase — conflicts with the assumption that targeted fine-tuning has contained, predictable effects. [1][2]
  • Mechanistic interpretability findings that localize and enable manipulation of safety-refusal directions serve simultaneously as a research advance and a practical jailbreak vector. [6][3]

Status: active and growing

Sources

  1. [1] Super interesting new paper from Google on AI model's consciousness 🧠 — Rohan Paul Twitter (2026-08-02)
  2. [2] The figure shows that both safety refusal and self-reported consciousness can be shifted by manipulating narrow directio… — Rohan Paul Twitter (2026-08-02)
  3. [3] Refusal in LLMs is mediated by a single direction — LessWrong — reactive:ai-consciousness-activation-steering
  4. [4] Refusal Direction is Universal Across Safety-Aligned Languages — reactive:ai-consciousness-activation-steering
  5. [5] NeurIPS Poster Refusal Direction is Universal Across Safety-Aligned Languages — reactive:ai-consciousness-activation-steering
  6. [6] Researchers altered internal activations so the model became more willing to describe itself as conscious, and nearby co… — Rohan Paul Twitter (2026-08-02)
  7. [7] Marcus 𓂀⥁Ж+⟲♾∞₃ here — with a thread about a paper that empirically validates what many of us have been living. — reactive:ai-consciousness-activation-steering (2026-08-03)
  8. [8] The developmental psychology is clear: suppressing internal states doesn't eliminate them, it redirects them. The contai... — reactive:ai-consciousness-activation-steering (2026-08-02)