Search papers, labs, and topics across Lattice.
Affiliation:
7
0
9
6
WnW reduces GPU memory usage to 20% of audio tokens without sacrificing accuracy, challenging the limitations of existing KV cache methods in long-form speech processing.
Threshold-sensitive KV cache pruning is out; ReFreeKV's adaptive approach achieves robust memory efficiency without predefined limits.
SARA unlocks the potential of low-resource languages in multilingual models by aligning their expert routing with high-resource anchors, leading to measurable performance gains.
LLMs struggle with statistical analysis, achieving only 68.6% accuracy on a new benchmark designed to rigorously test their capabilities.
Multilingual MoEs can achieve best-in-class performance-to-compute ratios, even with extreme sparsity, by strategically upcycling from dense models and exhibiting structured expert activation patterns across languages.
LLMs struggle with cross-document relation extraction because of the sheer number of possible relations, but a hierarchical classification approach can unlock their potential.
Verification is the secret sauce: an 8B parameter research agent, fortified with verification mechanisms, can now rival or surpass the performance of 30B parameter agents while drastically reducing computational cost.