Search papers, labs, and topics across Lattice.
This research introduces the "Compression Trinity," a novel framework that synergistically combines sparsity, quantization, and low-rank approximations to enhance the efficiency of Large Language Models (LLMs). By applying this unified approach to both the optimizer and model architecture, the study achieves significant improvements in convergence speed and accuracy, including up to 1.85x acceleration in pretraining and a 5.66% accuracy boost over state-of-the-art methods. The findings highlight that leveraging these three techniques together is crucial for overcoming the limitations of traditional compression methods, ultimately enabling more scalable and high-performance LLMs.
Jointly applying sparsity, quantization, and low-rank approximations can yield up to 5.66% better accuracy than the best existing methods for LLMs.
Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an accuracy-efficiency wall. This thesis proposes the"Compression Trinity,"a unified framework that applies the three pillars jointly: sparsity to reduce computation, quantization to minimize memory bandwidth, and low-rank approximations to recover accuracy. To accelerate pretraining, we apply the Trinity to the optimizer and model architecture. MKOR approximates curvature via block-diagonal sparsity and low-rank inversion, maintaining numerical stability for quantized states; it reduces curvature update complexity from $O(d^3)$ to $O(d^2)$ and accelerates convergence by up to 1.85x over KFAC. SLoPe accelerates training by up to 1.25x via a double-pruned backward pass for N:M sparsity, using low-rank"lazy"adapters in the final 1% of training to recover accuracy. For post-training compression, OPTIMA stabilizes static masks in a zero-training regime by formulating weight reconstruction as globally optimal column-wise quadratic programs, improving zero-shot accuracy by up to 3.97%. Given a fine-tuning budget, PATCH breaks the ceiling of static masks by learning a dynamic hybrid sparsity ratio between 0% and 50%, yielding up to 1.38x speedups. Finally, SLiM realizes the full Compression Trinity in one shot, using mathematically derived low-rank adapters to recover information lost to quantization and sparsity, improving accuracy by up to 5.66% over state-of-the-art methods and outperforming uncompressed dense models at equal parameter budgets by 0.6%. Together, these results show that jointly applying the Compression Trinity is essential for efficient, scalable, high-performance LLMs.