AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)
Alignment Forum · Rohin Shah · 2026-07-31
Google DeepMind's AGI Safety and Alignment Team publishes a two-year research retrospective covering chain-of-thought monitorability, frontier safety framework updates, deep alignment work on production Gemini models, interpretability pivots away from sparse autoencoders, amplified oversight via debate, and honeypot-based alignment evaluations.
Appears in
Extraction
Topics: agi-safetychain-of-thought-monitorabilityalignment-researchfrontier-safetyinterpretability
Claims
- GDM's work on chain-of-thought monitorability shifted field consensus toward viewing CoT as a valuable and preservable safety tool, particularly for difficult reasoning tasks where it is necessary for task completion.
- GDM was the first AI company to add a misalignment section to its Frontier Safety Framework, and the FSF now spans a cross-functional effort across many Google teams.
- The deep alignment team shifted focus to aligning current models on the grounds that they are similar enough to near-future systems that could accelerate AI research, making present-day alignment work likely to transfer.
- After finding primarily negative results applying sparse autoencoders to downstream safety tasks, GDM's interpretability team pivoted to more pragmatic approaches centered on probes, model forensics, and model diffing agents.
- Honeypot evaluations can detect early-stage schemers in realistic deployment settings, providing evidence about AI risk even if they would not catch highly competent schemers.
Key quotes
We are now fully in the midgame, and focus more on landing things in production.
while the chain of thought is frequently unfaithful, this tends to be on easy tasks where the chain of thought isn't load bearing to get the correct answer. However, this doesn't apply to difficult tasks requiring significant reasoning, as would be the case for many of the most concerning misalignment threat models.
Alignment is the only potentially scalable approach that we know of, and we want to make sure we don't lose sight of that goal.