The Information Machine

Value Generalisation 3: Pre-aligned AIs

Alignment Forum · Stuart_Armstrong · 2026-07-29

AI alignment researcher Stuart Armstrong proposes a 'pre-aligned AI' architecture that binds moral concepts to empirical learning so an AI's ethical constraints strengthen alongside its capabilities, structurally disadvantaging bad actors who must keep their AIs crippled to prevent the AI from recognizing misuse.

Open original ↗

Appears in

Extraction

Topics: ai-alignmentvalue-generalisationpre-aligned-aialignment-researchai-safety

Claims

  • A pre-aligned AI binds moral concepts to empirical concepts during learning, ensuring that increased capability automatically produces stronger ethical alignment rather than enabling misalignment.
  • Bad actors using a pre-aligned AI would be forced to keep it crippled to prevent it from gaining enough situational awareness to recognize harmful tasks, limiting its usefulness to them.
  • Initial morality for pre-aligned AIs consists of commercial, legal, and ethical components derived from examples and meta-ethical principles, without requiring fully rigorous or consistent data.
  • Architectural binding of moral and empirical concepts must be enforced so users cannot excise the moral component from the AI's code without breaking its empirical capabilities.
  • Pre-aligned AIs invert the standard alignment disadvantage: ethical users gain full AI capability by letting it learn freely, while bad actors must deliberately restrict theirs.

Key quotes

A pre-aligned AI is an AI whose morality increases with its capabilities.
The cost will be that they can't force the AI to behave unethically; the benefit will be that the full power of a generalising AI will be on their side, within those ethical bounds.
So, future users of these AIs, YOLO your way to power and alignment! The less you restrain them, the more they learn – and the more they learn, the more moral they become.