@GPU_MODE They focused on optimizing W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA de…
SemiAnalysis Twitter · SemiAnalysis (@SemiAnalysis_) · 2026-07-30
SemiAnalysis details the specific kernel optimization targets from the GPU_MODE AMD hackathon, covering W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA decode metadata planner on AMD MI355X hardware.
Appears in
Extraction
Topics: gpu-kernel-optimizationamd-mi355xmixture-of-expertstensor-parallelism
Claims
- The GPU_MODE hackathon team optimized W4A4 MoE kernels for AMD MI355X hardware.
- Top-K kernels and tensor-parallel all-reduce kernels were among the targeted optimizations.
- The MLA decode metadata planner was specifically improved as part of the hackathon work.
Key quotes
They focused on optimizing W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA decode metadata planner.