CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization
Abstract
Diffusion Transformers (DiTs) have achieved state-of-the-art(SOTA) performance in visual generative modeling, yet their training re-mains computationally prohibitive. While the recently proposed Momen-tum Orthogonalization (Muon) optimizer offers a promising alternativeto AdamW, its direct application to DiTs yields suboptimal late-stageconvergence. In this paper, we identify the root cause of this bottleneck:standard DiT architectures fuse functionally distinct weights (e.g., withinAdaLN and QKV layers) into unified tensors for computational efficiency.Applying Muon to these fused tensors inadvertently induces implicit sub-space coupling, which distorts update directions and degrades global op-timization. To address this, we introduce Chunked Muon (CMuon), asimple yet highly effective strategy that partitions these matrices intoindependent sub-components prior to orthogonalization. Extensive ex-periments demonstrate that a 675M-parameter DiT trained with CMuonachieves a FID of 1.18 on ImageNet 256 in just 200 epochs. This rep-resents more than a 2x training speedup over AdamW, while effectivelyovercoming the late-stage convergence plateaus of vanilla Muon.