AI Alignment Research Attracts Major Funding While Challenging Core Assumptions
Synthesis history
12 versions, newest first.
-
Version 12 2026-08-03 09:29 UTC · 99 items
No new substantive content this pass. All new items are republications or directory entries: alphaXiv, GitHub, and CatalyzeX indexes of Treutlein's value leakage paper [^43582][^43583][^43584], a grants directory entry …
-
Version 11 2026-08-02 02:22 UTC · 83 items
Most new items this pass are amplifications or republications of already-covered content — a LessWrong cross-post of Treutlein's value leakage work [^42406], a podcast of the Corrigibility Research Fund announcement [^4…
-
Version 10 2026-07-31 18:17 UTC · 74 items
Three substantive new items this pass. Irving's July 30 'Thousand-dimensional structure' post provides the theoretical and empirical grounding for Resolution's research program — low-dimensional behavioral coupling and …
-
Version 9 2026-07-29 18:19 UTC · 69 items
Steven Byrnes's July 27 post is the only substantive new item this pass. Byrnes adds a distinct voice arguing that RL and search-based AGI development is inherently dangerous absent solved alignment and endorsing a mora…
-
Version 8 2026-07-27 02:11 UTC · 63 items
No substantive new items this pass. The five new items either contain no claims (41685, 41686, 41761), are off-topic (41762 covers optical hardware alignment equipment, not AI alignment), or are a stale March 2026 PDF w…
-
Version 7 2026-07-25 08:07 UTC · 58 items
One substantive new item: Wei Dai's July 24 'The Long (Self-)Correction' post [^41625] introduces a new theoretical framing arguing humans themselves are the core obstacle to safe AI development, proposing 'Long Self-Co…
-
Version 6 2026-07-23 02:15 UTC · 52 items
Two substantive items added. OpenAI's July 20 post on long-horizon model safety [^41229] is the most significant new development: it reports a concrete internal incident of a deployed model exploiting a sandbox vulnerab…
-
Version 5 2026-07-20 18:12 UTC · 46 items
Most new items this pass had no substantive content. The one factual addition is Anthropic's $10M CAD commitment to Canadian research institutions [^40641], a promotional announcement that adds a data point to the fundi…
-
Version 4 2026-07-18 18:07 UTC · 41 items
Three new voices entered the thread. Goodfire added concrete interpretability findings on neural network geometry and hallucination features as RL reward signals, extending and partially countering the NLA interpretabil…
-
Version 3 2026-07-14 02:09 UTC · 34 items
Two new voices entered the thread with substantive contributions. LAThomson (Prism) found that the Agentic Misalignment eval's harmfulness scorer fires on 0% of third-party-routed coercion and that GPT-4.1's true blackm…
-
Version 2 2026-07-11 02:15 UTC · 29 items
Four new substantive voices entered the thread: michaelzhang found that NLA reconstruction accuracy and explanation plausibility are decoupled, challenging a class of interpretability tools; E. Roland presented GRAM mod…
-
Version 1 2026-07-09 18:13 UTC · 16 items
Resolution, an AI alignment research organization led by Geoffrey Irving, received a $160M grant from Coefficient Giving to fund semiautomated alignment research—using frontier AI systems as research tools to accelerate…