Search papers, labs, and topics across Lattice.
5
0
7
2
LSA slashes indexing overhead while maintaining full attention performance, enabling efficient long-context processing for models with up to one million tokens.
CoSA achieves nearly 5脳 faster attention computation without sacrificing accuracy, revolutionizing long-context processing in LLMs.
AngelSpec achieves nearly double the inference speed of traditional methods while improving output quality by intelligently adapting to the specific demands of different tasks.
PIVOT achieves up to 4x faster indexing for token-level sparse attention without sacrificing accuracy, transforming how we handle query processing in large models.
D-Cut transforms speculative decoding efficiency by cutting verification costs, achieving up to 3.0x speedup over traditional methods in high-concurrency scenarios.