Search papers, labs, and topics across Lattice.
6
0
6
8
Pre-attention spikes and inter-spike plateaus reveal a surprising organization in hybrid linear attention models that could redefine our understanding of activation dynamics in LLMs.
Scaling laws for contrastive learning reveal that learning interactions between views fundamentally alters optimization dynamics compared to linear regression.
Achieving up to 88x efficiency gains, Taylor-Calibrate transforms the way we initialize hybrid linear attention models, drastically reducing the training burden.
Forget fancy quantization schemes – a simple token-wise INT4 quantization with Hadamard rotation is all you need to nearly match FP16 accuracy in LLM serving, without sacrificing throughput.
Diffusion language models can now match autoregressive quality, thanks to a clever trick that forces them to agree with themselves.
Verifier-free evolution can now match or exceed the performance of verifier-based methods, while slashing API costs by 3x and boosting throughput by 10x, thanks to a clever model orchestration strategy.