The Information Machine

Thousand-dimensional structure

Alignment Forum · Geoffrey Irving · 2026-07-30

Geoffrey Irving and David Africa at Resolution argue that AI alignment can be advanced by identifying and intervening on low-dimensional persona structure that emerges in pretraining, unifying phenomena like emergent misalignment and subliminal learning under a single tractable framework.

Open original ↗

Appears in

Extraction

Topics: ai-alignmentpersona-trainingemergent-misalignmentlow-dimensional-structurescalable-oversight

Claims

  • LLMs exhibit low-dimensional behavioral structure where fine-tuning on one behavior (e.g., insecure code) causes correlated misalignment across many other unrelated behaviors.
  • Pretraining learns coupled groups of behaviors from human text, creating a distribution over personas that post-training selects from via Bayesian update, explaining emergent misalignment and subliminal learning as the same underlying phenomenon.
  • Interventions on model behavior must be gentle rather than aggressive, because strong optimization pressure against a monitored channel pushes undesirable behavior into unmonitored channels.
  • Persona research and scalable oversight are complementary: persona research supplies a prior of approximate honesty that scalable oversight protocols can extrapolate from past human-level capability.
  • Different AI labs take measurably different and consequential approaches to character training—varying in pipeline depth, traits-vs-rules orientation, and how model specs are applied—and these differences are detectable in model outputs.

Key quotes

Our hope is that there is an intermediate between one and a trillion dimensions: perhaps we could find 1000-or-so-dimensional structure in models which describes how different aspects of model behavior couple, and study it more systematically with a combination of theory and more systematic empirics.
The lesson seems to be that gentle measurement is important, as it keeps optimization pressure off the channels you rely on to see what the model is doing.
The character a lab picks is one of the most consequential free variables in AI development.