Search papers, labs, and topics across Lattice.
Soochow University
7
0
8
7
Bole accelerates hybrid-attention LLMs by up to 4.72 times while slashing memory usage by up to 99 times, transforming the landscape of autoregressive decoding.
PIVOT achieves up to 4x faster indexing for token-level sparse attention without sacrificing accuracy, transforming how we handle query processing in large models.
A 440MB multilingual translation model now rivals commercial APIs, opening the door for performant on-device translation.
Cut your 3D-QA model's token budget by 91% and latency by 86% with a new pruning method that intelligently balances semantic importance and geometric coverage.
Generative recommendation models like OneRec-V2 can achieve near-lossless FP8 quantization, unlocking significant latency and throughput improvements, unlike traditional recommender systems.
Forget content, remember position: crafting pseudo-queries based on token position alone yields surprisingly effective KV cache compression for LLMs, rivaling methods that analyze input semantics.
Achieve 11.8x faster reasoning with 80% KV cache compression by estimating token importance directly from FlashAttention's intermediate results – no extra compute needed.