The Information Machine

AI Alignment Research Attracts Major Funding While Challenging Core Assumptions

Synthesis history

12 versions, newest first.

  1. Version 12 2026-08-03 09:29 UTC · 99 items

    No new substantive content this pass. All new items are republications or directory entries: alphaXiv, GitHub, and CatalyzeX indexes of Treutlein's value leakage paper [^43582][^43583][^43584], a grants directory entry …

  2. Version 11 2026-08-02 02:22 UTC · 83 items

    Most new items this pass are amplifications or republications of already-covered content — a LessWrong cross-post of Treutlein's value leakage work [^42406], a podcast of the Corrigibility Research Fund announcement [^4…

  3. Version 10 2026-07-31 18:17 UTC · 74 items

    Three substantive new items this pass. Irving's July 30 'Thousand-dimensional structure' post provides the theoretical and empirical grounding for Resolution's research program — low-dimensional behavioral coupling and …

  4. Version 9 2026-07-29 18:19 UTC · 69 items

    Steven Byrnes's July 27 post is the only substantive new item this pass. Byrnes adds a distinct voice arguing that RL and search-based AGI development is inherently dangerous absent solved alignment and endorsing a mora…

  5. Version 8 2026-07-27 02:11 UTC · 63 items

    No substantive new items this pass. The five new items either contain no claims (41685, 41686, 41761), are off-topic (41762 covers optical hardware alignment equipment, not AI alignment), or are a stale March 2026 PDF w…

  6. Version 7 2026-07-25 08:07 UTC · 58 items

    One substantive new item: Wei Dai's July 24 'The Long (Self-)Correction' post [^41625] introduces a new theoretical framing arguing humans themselves are the core obstacle to safe AI development, proposing 'Long Self-Co…

  7. Version 6 2026-07-23 02:15 UTC · 52 items

    Two substantive items added. OpenAI's July 20 post on long-horizon model safety [^41229] is the most significant new development: it reports a concrete internal incident of a deployed model exploiting a sandbox vulnerab…

  8. Version 5 2026-07-20 18:12 UTC · 46 items

    Most new items this pass had no substantive content. The one factual addition is Anthropic's $10M CAD commitment to Canadian research institutions [^40641], a promotional announcement that adds a data point to the fundi…

  9. Version 4 2026-07-18 18:07 UTC · 41 items

    Three new voices entered the thread. Goodfire added concrete interpretability findings on neural network geometry and hallucination features as RL reward signals, extending and partially countering the NLA interpretabil…

  10. Version 3 2026-07-14 02:09 UTC · 34 items

    Two new voices entered the thread with substantive contributions. LAThomson (Prism) found that the Agentic Misalignment eval's harmfulness scorer fires on 0% of third-party-routed coercion and that GPT-4.1's true blackm…

  11. Version 2 2026-07-11 02:15 UTC · 29 items

    Four new substantive voices entered the thread: michaelzhang found that NLA reconstruction accuracy and explanation plausibility are decoupled, challenging a class of interpretability tools; E. Roland presented GRAM mod…

  12. Version 1 2026-07-09 18:13 UTC · 16 items

    Resolution, an AI alignment research organization led by Geoffrey Irving, received a $160M grant from Coefficient Giving to fund semiautomated alignment research—using frontier AI systems as research tools to accelerate…