Search papers, labs, and topics across Lattice.
3
0
7
1
Achieving 2-3x near-lossless compression of KV caches could revolutionize the efficiency of transformer inference without sacrificing performance.
LLM development is flying blind by ignoring causal inference, leaving models vulnerable to confounding and distribution shifts throughout pretraining, alignment, and evaluation.
Masked diffusion language models can now achieve 21.8x better compute efficiency than autoregressive models, thanks to binary encoding and index shuffling.