The Information Machine

@GPU_MODE They focused on optimizing W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA de…

SemiAnalysis Twitter · SemiAnalysis (@SemiAnalysis_) · 2026-07-30

SemiAnalysis details the specific kernel optimization targets from the GPU_MODE AMD hackathon, covering W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA decode metadata planner on AMD MI355X hardware.

Open original ↗

Appears in

Extraction

Topics: gpu-kernel-optimizationamd-mi355xmixture-of-expertstensor-parallelism

Claims

  • The GPU_MODE hackathon team optimized W4A4 MoE kernels for AMD MI355X hardware.
  • Top-K kernels and tensor-parallel all-reduce kernels were among the targeted optimizations.
  • The MLA decode metadata planner was specifically improved as part of the hackathon work.

Key quotes

They focused on optimizing W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA decode metadata planner.