Search papers, labs, and topics across Lattice.
Affiliation:
1
0
1
1
Jointly pruning and quantizing weights under a single Bayesian objective breaks the traditional trade-offs of sequential compression pipelines, squeezing modern LLMs like Llama 3.2 and Qwen 2.5 further without catastrophic accuracy degradation.