Search papers, labs, and topics across Lattice.
2
0
4
3
Enhancing on-policy reinforcement learning can paradoxically reduce the diversity of successful behaviors, leading to a trade-off that challenges future trainability.
Turns out, telling LLMs *not* to use the answer when generating reverse chain-of-thought reasoning can actually make them *more* reliant on it鈥攂ut a skeleton-guided approach breaks the cycle.