Search papers, labs, and topics across Lattice.
2
0
4
7
Cooperative multi-agent training can unlock unsupervised reasoning capabilities in RL, yielding performance gains that rival supervised methods without the need for costly annotations.
Freezing most of your critic network and only training a tiny LoRA adapter can dramatically improve off-policy RL performance and stability.