Search papers, labs, and topics across Lattice.
5
0
8
5
HiLS-Attention achieves over 64x context length extrapolation with 90% retrieval accuracy, outperforming traditional full attention mechanisms.
TACO reveals that agentic models can learn to optimize tool usage without external judges, achieving higher accuracy and efficiency in multimodal tasks.
LLMs can maintain long-context performance even with aggressive KV-cache eviction by learning to predict token importance and compressing evicted tokens into a latent memory.
Quantizing rollouts in LLM RL pipelines introduces a training-inference gap that QaRL closes, leading to +5.5 performance on math problems.
Achieve near-lossless 2-bit LLMs with a novel quantization-aware training scheme that progressively reduces precision and intelligently handles outlier channels.