Search papers, labs, and topics across Lattice.
This paper introduces Interleaved Cross-Block Quantization (ICBQ), a novel scheduling technique that revisits the boundary between consecutive chunks during block-wise post-training quantization of large language models. By refining seam pairs twice, ICBQ effectively reduces errors that accumulate in traditional sequential methods, leading to improved performance in quantized models. Experimental results show that ICBQ significantly lowers perplexity compared to the baseline, demonstrating its robustness even in configurations where the baseline fails.
ICBQ not only cuts perplexity in quantized models but also salvages performance where traditional methods falter, redefining efficiency in model compression.
Compressing large language models to two bits or fewer is increasingly feasible through block-wise post-training quantization; cross-block variants reconstruct neighboring Transformer blocks within a moving window. In the fixed two-block setting studied here, the matched sequential baseline moves this window through the network once, so errors introduced early in the sweep are not revisited. We propose Interleaved Cross-Block Quantization (ICBQ), a scheduling modification that revisits the boundary pair between consecutive chunks. Each seam pair is refined twice: first at the end of one chunk and again at the start of the next. The method retains the local two-block objective and reuses the calibration inputs of existing block-wise PTQ pipelines. Under stated local contraction and smoothness assumptions, we derive a depth-wise upper-bound comparison in which seam revisits multiply the propagated term while the residual remains bounded independently of depth. In the reported experiments, ICBQ reduces ternary-quantization perplexity relative to the matched Sequential CBQ baseline, yields finite perplexity in configurations where the baseline has severe degradation, and can also be used with 3-bit and 2-bit GPTQ.