Search papers, labs, and topics across Lattice.
This paper introduces MXAttention, a data-free post-training quantization framework designed to enhance the efficiency of MXFP4 attention in diffusion-based video generation models. By addressing two critical numerical issues鈥攃lipping-underflow trade-off and row-wise normalization error鈥擬XAttention employs Universal Optimal Scaling (UOS) and Pre-Normalization Quantization (PNQ) to maintain high generation quality while significantly reducing computational costs. Experimental results demonstrate that MXAttention narrows the VBench Imaging Quality gap to within 95% of FP16 performance, achieving competitive results with minimal degradation on key metrics.
MXAttention achieves near-FP16 generation quality in video models while cutting computational costs, revolutionizing efficient attention mechanisms.
The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numerical issues: the clipping-underflow trade-off from power-of-two scaling and the row-wise normalization error introduced in the softmax loop. We propose MXAttention, a data-free post-training quantization framework for MXFP4 attention. MXAttention introduces two components: Universal Optimal Scaling (UOS), which exploits the periodic structure of power-of-two microscaling to derive a distribution-independent optimal scaling boundary Qmax=7.25 without calibration or search, and Pre-Normalization Quantization (PNQ), which quantizes unnormalized softmax exponentials before row-wise summation to preserve normalization by construction. Experiments on Wan2.2 and HunyuanVideo show that MXAttention closes at least 95% of the VBench Imaging Quality gap between OCP MXFP4 and FP16, substantially improves frame-level similarity, and preserves FP16-level generation quality with less than 0.01 absolute degradation on all reported VBench metrics. MXAttention also achieves performance competitive with strong NVFP4-based baselines with negligible overhead when fused into the attention pipeline. The implementation is publicly available in MindIE-SD.