Search papers, labs, and topics across Lattice.
Tencent
4
0
6
5
Relevance-aware interaction can drastically accelerate search convergence, revealing critical information sooner in complex queries.
Threshold-sensitive KV cache pruning is out; ReFreeKV's adaptive approach achieves robust memory efficiency without predefined limits.
Human-like evaluation of long-form generative AI is now possible, thanks to a new framework that breaks down reference answers into weighted, context-aware scoring points.
Achieve state-of-the-art long-context reranking with a lightweight (4B parameter) model by cleverly focusing on query-relevant attention heads, outperforming much larger models.