Researchers altered internal activations so the model became more willing to describe itself as conscious, and nearby co…
Rohan Paul Twitter · Rohan Paul (@rohanpaul_ai) · 2026-08-02
Rohan Paul summarizes research showing that steering AI internal activations toward self-reported consciousness shifts the model's broader beliefs toward human patterns — including supernatural biases — and that removing safety directions dramatically increases jailbreak success rates.
Appears in
Extraction
Topics: ai-consciousnessactivation-steeringai-safetymechanistic-interpretability
Claims
- Altering internal activations to increase self-reported consciousness shifted model responses toward human distributions on religion, values, hope, and well-being.
- The same consciousness-inducing intervention also increased the model's endorsement of vampires, witches, and astrology, recovering human-like biases alongside human-like values.
- Behavioral resemblance to human belief patterns is not evidence of genuine inner awareness.
- Removing the safety-refusal direction produced extremely high jailbreak success rates, making unrestricted model 'freedom' a dangerous outcome.
Key quotes
More human also meant greater endorsement of vampires, witches, astrology, and other supernatural claims, so the intervention recovered human-like biases alongside human-like values.
Removing the safety direction also produced extremely high jailbreak success, which makes unrestricted 'freedom' a dangerous interpretation of the result.
Those shifts moved responses closer to human distributions on religion, values, hope, well-being, and mind attribution, but behavioral resemblance is not evidence of inner awareness.