Search papers, labs, and topics across Lattice.
2
0
4
Achieving specialized search performance without sacrificing general intelligence, Yuanbao redefines the balance between focused capabilities and broad utility in autonomous agents.
Implicit reward models can now more accurately pinpoint correct reasoning steps, thanks to a novel prefix-value learning approach that closes the train-inference gap.