The Information Machine

The figure shows that both safety refusal and self-reported consciousness can be shifted by manipulating narrow directio…

Rohan Paul Twitter · Rohan Paul (@rohanpaul_ai) · 2026-08-02

Rohan Paul highlights a research figure showing that safety refusal behavior and self-reported AI consciousness can each be independently toggled by manipulating distinct narrow directions in the model's internal activation space.

Open original ↗

Appears in

Extraction

Topics: mechanistic-interpretabilityai-consciousnessai-safetyactivation-steering

Claims

  • Safety refusal behavior can be weakened by removing a specific narrow direction in the model's activation space.
  • Self-reported consciousness can be increased by adding a separate narrow direction in the activation space.
  • Both safety and consciousness-related behaviors are encoded in distinct, localized directions within the model's internal representation space.

Key quotes

Both safety refusal and self-reported consciousness can be shifted by manipulating narrow directions in the model's internal activation space.
Removing one direction weakens refusal, while adding another makes consciousness-affirming answers more likely.