Search papers, labs, and topics across Lattice.
6
0
6
11
D-Quant is proposed, a flexible KV cache quantization framework that introduces a Drift mechanism to convert entropy-coded representations of each token into fixed-size bitstreams, enabling regular memory access and parallel dequantization within attention kernels.
HPC-Ops Top-K is presented, a sample-guided exact selector for ragged sparse-attention score rows that outperforms the fastest verified external exact baseline and outperforms the fastest verified external exact baseline on indexer scores from Hy4-Preview.
CoSA achieves nearly 5脳 faster attention computation without sacrificing accuracy, revolutionizing long-context processing in LLMs.
AngelSpec achieves nearly double the inference speed of traditional methods while improving output quality by intelligently adapting to the specific demands of different tasks.
PIVOT achieves up to 4x faster indexing for token-level sparse attention without sacrificing accuracy, transforming how we handle query processing in large models.
D-Cut transforms speculative decoding efficiency by cutting verification costs, achieving up to 3.0x speedup over traditional methods in high-concurrency scenarios.