The Information Machine

AI Alignment Research Revisits Filtering and Steering Interventions

Synthesis history

4 versions, newest first.

  1. Version 4 2026-06-25 18:18 UTC · 35 items

    Three substantive new items extend the thread beyond training interventions into monitoring and interpretability. Engels adds a DiffusionGemma transparency analysis showing monitorability parity with autoregressive Gemm…

  2. Version 3 2026-06-21 08:11 UTC · 26 items

    Item 31225 adds a fourth intervention type: OpenAI RL training on realistic human scenarios, which produces cross-domain safety transfer (health-only training improved non-health behaviors). This creates a new tension w…

  3. Version 2 2026-06-19 02:20 UTC · 21 items

    Item 29914 adds a third substantive thread: Google DeepMind's CallumMcDougall reports synthetic document and chat finetuning can produce OOD safety improvements on Gemini 3 Flash, with documented failure modes around su…

  4. Version 1 2026-06-15 08:07 UTC · 13 items

    Alignment researchers are producing empirical results that challenge two widely-used intervention approaches: sparse autoencoder (SAE) based model steering and supervised fine-tuning (SFT) data filtering for safety. A p…