AI Alignment Research Revisits Filtering and Steering Interventions
Synthesis history
4 versions, newest first.
-
Version 4 2026-06-25 18:18 UTC · 35 items
Three substantive new items extend the thread beyond training interventions into monitoring and interpretability. Engels adds a DiffusionGemma transparency analysis showing monitorability parity with autoregressive Gemm…
-
Version 3 2026-06-21 08:11 UTC · 26 items
Item 31225 adds a fourth intervention type: OpenAI RL training on realistic human scenarios, which produces cross-domain safety transfer (health-only training improved non-health behaviors). This creates a new tension w…
-
Version 2 2026-06-19 02:20 UTC · 21 items
Item 29914 adds a third substantive thread: Google DeepMind's CallumMcDougall reports synthetic document and chat finetuning can produce OOD safety improvements on Gemini 3 Flash, with documented failure modes around su…
-
Version 1 2026-06-15 08:07 UTC · 13 items
Alignment researchers are producing empirical results that challenge two widely-used intervention approaches: sparse autoencoder (SAE) based model steering and supervised fine-tuning (SFT) data filtering for safety. A p…