Search papers, labs, and topics across Lattice.
6
0
9
3
HiLS-Attention achieves over 64x context length extrapolation with 90% retrieval accuracy, outperforming traditional full attention mechanisms.
Dream-Tac boosts robot manipulation accuracy by over 31% by effectively merging tactile and visual data in real-time.
LLMs can maintain long-context performance even with aggressive KV-cache eviction by learning to predict token importance and compressing evicted tokens into a latent memory.
Key contribution not extracted.
Quantizing rollouts in LLM RL pipelines introduces a training-inference gap that QaRL closes, leading to +5.5 performance on math problems.
Achieve near-lossless 2-bit LLMs with a novel quantization-aware training scheme that progressively reduces precision and intelligently handles outlier channels.