The Long (Self-)Correction
Alignment Forum · Wei Dai · 2026-07-24
Alignment Forum researcher Wei Dai proposes 'Long Self-Correction' as a new conceptual frame for AI risk discourse, arguing that humans must fix deep philosophical, moral, and strategic flaws in themselves before they can safely build or oversee powerful AI systems.
Appears in
Extraction
Topics: ai-safetylong-reflectionai-governancehuman-valuesexistential-risk
Claims
- The concept of 'AI Pause' is underspecified because it does not address that humans themselves are unsafe builders and overseers of powerful AI.
- 'Long Reflection' incorrectly implies that additional thinking time is the primary bottleneck, rather than fundamental human flaws.
- Humans lack a workable moral framework, are poorly calibrated about their own philosophical competence, and are susceptible to manipulation via sycophancy and plausible-sounding arguments.
- Human morality functions in practice as a status game that actively discourages careful strategy and philosophy.
- Humans have nonetheless made slow, mysterious progress on these issues over very long timescales, which is the basis for cautious optimism about Long Self-Correction eventually succeeding.
Key quotes
humans aren't ready to build AIs or other extremely powerful technologies, because we're currently too flawed, in a variety of ways, and it will take a long process (which may or may not end up succeeding) to fix those flaws.
My main hope for a Long Self-Correction eventually succeeding rests on the fact that humans have seemingly, mysteriously, made progress on these issues over a very long period of time.
in practice, human morality is a kind of status game that actively disvalues careful strategy and philosophy in most places