Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
0
Complex scoring algorithms for KV cache compression are largely unnecessary: protecting the initial prompt and dropping reasoning tokens completely at random matches state-of-the-art accuracy with up to 43% higher serving throughput.
Test-time harnesses can nearly double the performance of weaker models, transforming how we think about capability transfer in AI.
Whisper's speech-centric training leaves audio-LLMs tone-deaf to music and environmental sounds, but a simple fine-tune can fix that.