Search papers, labs, and topics across Lattice.
6
0
8
4
D-Quant is proposed, a flexible KV cache quantization framework that introduces a Drift mechanism to convert entropy-coded representations of each token into fixed-size bitstreams, enabling regular memory access and parallel dequantization within attention kernels.
CoSA achieves nearly 5脳 faster attention computation without sacrificing accuracy, revolutionizing long-context processing in LLMs.
AngelSpec achieves nearly double the inference speed of traditional methods while improving output quality by intelligently adapting to the specific demands of different tasks.
PIVOT achieves up to 4x faster indexing for token-level sparse attention without sacrificing accuracy, transforming how we handle query processing in large models.
D-Cut transforms speculative decoding efficiency by cutting verification costs, achieving up to 3.0x speedup over traditional methods in high-concurrency scenarios.
Achieving a 6.37x speedup in inference while expanding OCR capabilities across long-tail tasks sets a new benchmark for lightweight models.