Search papers, labs, and topics across Lattice.
Affiliation:
4
0
7
0
Complex scoring algorithms for KV cache compression are largely unnecessary: protecting the initial prompt and dropping reasoning tokens completely at random matches state-of-the-art accuracy with up to 43% higher serving throughput.
Test-time harnesses can nearly double the performance of weaker models, transforming how we think about capability transfer in AI.
Whisper's speech-centric training leaves audio-LLMs tone-deaf to music and environmental sounds, but a simple fine-tune can fix that.
Editing a rule in an LLM is not like editing a fact; you can't just tweak one layer – formulas live in early layers, instances in the middle.