Search papers, labs, and topics across Lattice.
4
0
6
2
Pre-attention spikes and inter-spike plateaus reveal a surprising organization in hybrid linear attention models that could redefine our understanding of activation dynamics in LLMs.
Attention Sink, where Transformers fixate on seemingly irrelevant tokens, is more than just a quirk – it's a fundamental challenge impacting training, inference, and even causing hallucinations, demanding a systematic approach to understanding and mitigating its effects.
Achieve better compression in low-bit quantization by considering not just numerical sensitivity, but also the structural role of each layer.
Forget fixed uniform intervals: BPDQ unlocks high-fidelity 2-bit quantization for LLMs by adaptively shaping the quantization grid.