Value Generalisation 2: The Missing Hole in AIs’ abilities
Alignment Forum · Stuart_Armstrong · 2026-07-29
Alignment researcher Stuart Armstrong argues that LLMs lack a fundamental human cognitive capability he calls 'strong generalisation,' which explains why they perpetually appear near-AGI without reaching it and why reliable AI value alignment cannot be achieved through scaling, more data, or chain-of-thought techniques alone.
Appears in
Extraction
Topics: llm-limitationsstrong-generalisationagivalue-generalisationalignment-research
Claims
- LLMs lack 'strong generalisation,' a cluster of human abilities including situational awareness, out-of-distribution generalization, adaptive world modeling, and long-range planning.
- LLMs succeed on long-horizon benchmarks through pattern matching rather than genuine generalisation, which is why benchmark saturation does not translate to real-world AGI performance.
- The absence of strong generalisation explains why subject matter experts get far more out of LLMs than amateurs—experts supply the generalisation capability the LLM lacks.
- Value generalisation is fundamentally harder than empirical generalisation because values have no ground truth and morally important concepts cannot be discarded the way empirical ones can.
- Current scaling approaches and reasoning techniques like chain-of-thought do not address the missing strong generalisation ability and are therefore unlikely to produce reliable alignment.
Key quotes
I think that humans have an ability, or a mix of abilities, that I'm calling strong generalisation. And LLMs and their derived models lack strong generalisation.
Since we can't define strong generalisation or its various components (like long-term planning) the best we can do is produce examples (benchmarks) that illustrate it. And LLMs will succeed on these benchmarks using the simplest pattern they can find, which is not strong generalisation.
Values, goals, preferences, objectives: these are similar concepts in that their ground truths are not empirically defined, and we try to capture them imperfectly in formal definitions or training data, neither of which are enough to pin down these concepts across all environments.