The Information Machine

Value Generalisation 1: a Research and Deployment Program

Alignment Forum · Stuart_Armstrong · 2026-07-29

Stuart Armstrong proposes a three-stage research and commercial deployment program on value generalisation — the ability for AI to correctly extend human values to novel situations — arguing it is a specific missing capability necessary for trustworthy autonomous AI agents.

Open original ↗

Appears in

Extraction

Topics: value-generalisationai-alignmentautonomous-agentsout-of-distribution-generalizationai-safety

Claims

  • Current AI systems cannot reliably extrapolate human values to novel situations, making them untrustworthy for autonomous delegation even though the raw capabilities for such delegation already exist.
  • Value generalisation is a specific missing capability that scaling and patching alone will not deliver.
  • Capabilities can advance without explicit value generalisation, but alignment cannot — making early, deliberate development of this capability critical.
  • Commercial deployment of alignment techniques is preferable to academic publication because papers get ignored or mined for capability-relevant parts while alignment components are discarded.
  • A three-stage program — out-of-distribution recognition, relevant concept selection, and full explicit value generalisation — can progressively build toward pre-aligned AIs whose alignment grows with their capabilities.

Key quotes

today's AIs cannot be relied on to understand your interests in situations that weren't covered – explicitly or implicitly – by their training and instructions. They extrapolate patterns naively. Push them past the situations they were shaped for, and they will still confidently do something; it just won't reliably be what you wanted.
capabilities don't need explicit generalisation; alignment does. AIs may become very powerful without ever developing explicit generalisation – many companies are effectively betting on exactly that.
The way to make value generalisation matter is to make it one of the most useful things on the market: learning AIs that can actually be trusted with delegation, deployed widely, with the alignment machinery the load-bearing piece that makes them reliable and trustworthy.