Search papers, labs, and topics across Lattice.
4
0
8
4
Coding agents can significantly improve payment integration performance with targeted skills, achieving up to 91.37% success in complex scenarios.
OPRD closes the performance gap between student and teacher models while training 1.44x faster and using 54% less memory than traditional methods.
Achieve state-of-the-art results in agentic knowledge base question answering by distilling gold-action policies into on-policy student rollouts, bridging the gap between sparse rewards and weakly supervised intermediate actions.
RLVR models exhibit "Early Correctness Coherence" under noisy supervision, suggesting a surprising opportunity for self-correction via dynamic label refinement.