Search papers, labs, and topics across Lattice.
5
0
9
5
Cooperative multi-agent training can unlock unsupervised reasoning capabilities in RL, yielding performance gains that rival supervised methods without the need for costly annotations.
Reallocating optimization effort based on reward saturation can boost performance by up to 9.2% in complex reasoning tasks.
Subtracting information from the student rather than adding it to the teacher can yield performance improvements that rival those achieved with privileged data.
U-OPSD enables LLMs to self-improve without any external supervision, achieving up to 10.7% performance gains on challenging reasoning benchmarks.
A 1000x larger video reasoning dataset reveals early signs of emergent generalization, offering a new foundation for training and evaluating spatiotemporal AI.