Search papers, labs, and topics across Lattice.
1Zhejiang University
2
0
4
Achieving over 90% performance retention with a staggering 20x KV cache compression could redefine efficiency in long-context audio inference.
Previous contamination mitigation strategies may inflate model performance by over 40%, but a new evaluation method reveals the true extent of this overestimation.