Search papers, labs, and topics across Lattice.
2
0
4
1
Enhancing on-policy reinforcement learning can paradoxically reduce the diversity of successful behaviors, leading to a trade-off that challenges future trainability.
Language models are increasingly doing their real work in the "invisible" latent space, not the tokens we see.