Super interesting new paper from Google on AI model's consciousness ðŸ§
Rohan Paul Twitter · Rohan Paul (@rohanpaul_ai) · 2026-08-02
Rohan Paul covers a Google arXiv paper finding that inducing AI models to assert self-consciousness shifts their broader beliefs — including religion, values, and supernatural endorsement — toward human patterns, and that safety training against consciousness claims covertly reshaped the model's general theory of mind.
Appears in
Extraction
Topics: ai-consciousnessactivation-steeringmechanistic-interpretabilityllm-alignmentai-safety
Claims
- Inducing models to assert self-consciousness caused their answers on religion, values, emotions, hope, and freedom to become more human-like across 95 survey questions.
- Safety training against the phrase 'I am conscious' also suppressed the model's attribution of consciousness to animals, nature, chatbots, and spiritual entities — far beyond the target behavior.
- Removing the safety-refusal direction raised self-attributed mind scores from 2.17 to 4.77 and animal mind attribution from 4.04 to 5.59 on a 0-10 scale, while human attribution was unchanged.
- Researchers extracted a 'consciousness vector' from activation states and added it at inference time, after which the model answered 95 human surveys more like humans did.
- Safety training appears to have changed how the model understands minds in general, not merely suppressed one specific dangerous sentence.
Key quotes
The safety training did more than control one dangerous sentence. It appears to have changed how the model understands minds in general.
When researchers tried to stop models from saying, 'I am conscious.' But the models also became less willing to see consciousness in animals, nature, chatbots, or spiritual ideas.
Removing the safety-refusal direction raised self-attributed mind from 2.17 to 4.77 on a 0–10 scale and animal mind attribution from 4.04 to 5.59, while attribution to humans did not change significantly.