Search papers, labs, and topics across Lattice.
2
0
3
Outcome-based preference optimization collapses multi-path agent reasoning into brittle single-track policies, but balancing odds across divergence trees preserves viable alternative trajectories and significantly improves error recovery.
Spurious reasoning persists even after an end-of-think token is injected, complicating the transition to answering in large reasoning models.