Search papers, labs, and topics across Lattice.
Zhejiang University
2
0
4
Dense retrievers miss the mark on answerability, dropping QA performance drastically despite high semantic relevance, revealing a critical oversight in RAG systems.
ISPO reduces critical reasoning failures in RLVR by transforming reward structures, leading to superior performance on complex reasoning tasks.