Search papers, labs, and topics across Lattice.
This paper introduces SwiftQK, a novel multi-GPU RMSNorm kernel that optimizes Query-Key Normalization (QK-Norm) by minimizing cross-GPU communication. By exchanging only scalar normalization statistics and overlapping Peer-to-Peer reductions with independent computations, SwiftQK achieves significant reductions in latency. Evaluations demonstrate that SwiftQK can decrease QK-Norm latency by up to 93.9% and improve end-to-end serving efficiency by 29.5% compared to traditional methods.
SwiftQK slashes QK-Norm latency by up to 93.9%, revolutionizing multi-GPU training efficiency for large language models.
Query-Key Normalization (QK-Norm) improves the training stability and quality of modern Large Language Models (LLMs). However, under Tensor Parallelism (TP), layerwise QK-Norm introduces additional cross-GPU communication because the normalization factor depends on the full hidden vector. We present SwiftQK, a multi-GPU RMSNorm kernel that exchanges only scalar normalization statistics and overlaps the remaining Peer-to-Peer reduction with independent element-wise computation in a deadlock-safe persistent kernel. Evaluations on recent LLMs show that SwiftQK reduces QK-Norm latency by 81.4--93.9% relative to the standard TP QK-Norm using full-vector All-Gather. In end-to-end serving, SwiftQK reduces TPOT on average by 29.5% over the All-Gather-based baseline and by 14.3% over an optimized scalar-aggregation implementation.