Search papers, labs, and topics across Lattice.
Ant Group
4
0
6
Despite high rationale identification, search-augmented models struggle with refusal, achieving only 42.9% correct halting on unanswerable multi-hop questions.
EviSD achieves state-of-the-art performance in question-answering tasks by leveraging privileged evidence, outperforming existing methods while maintaining efficiency in response generation.
OPRD closes the performance gap between student and teacher models while training 1.44x faster and using 54% less memory than traditional methods.
Achieve state-of-the-art results in agentic knowledge base question answering by distilling gold-action policies into on-policy student rollouts, bridging the gap between sparse rewards and weakly supervised intermediate actions.