The Information Machine

Transformer Attention: A Decade of Innovation Recognized by SemiAnalysis

Synthesis history

2 versions, newest first.

  1. Version 2 2026-06-30 08:23 UTC · 28 items

    The substantive addition this pass is NVIDIA cuDNN's official DSA documentation [^36045], which gives sparse attention a concrete production infrastructure path and establishes NVIDIA as a new perspective voice active o…

  2. Version 1 2026-06-29 02:17 UTC · 22 items

    SemiAnalysis published a community recognition thread on June 29, 2026, tracing transformer attention from the 2017 Multi-Head Attention paper through FlashAttention, PagedAttention/vLLM, and the current wave of linear …