The figure shows that both safety refusal and self-reported consciousness can be shifted by manipulating narrow directio…
Rohan Paul Twitter · Rohan Paul (@rohanpaul_ai) · 2026-08-02
Rohan Paul highlights a research figure showing that safety refusal behavior and self-reported AI consciousness can each be independently toggled by manipulating distinct narrow directions in the model's internal activation space.
Appears in
Extraction
Topics: mechanistic-interpretabilityai-consciousnessai-safetyactivation-steering
Claims
- Safety refusal behavior can be weakened by removing a specific narrow direction in the model's activation space.
- Self-reported consciousness can be increased by adding a separate narrow direction in the activation space.
- Both safety and consciousness-related behaviors are encoded in distinct, localized directions within the model's internal representation space.
Key quotes
Both safety refusal and self-reported consciousness can be shifted by manipulating narrow directions in the model's internal activation space.
Removing one direction weakens refusal, while adding another makes consciousness-affirming answers more likely.