Search papers, labs, and topics across Lattice.
4
0
5
3
Cooperative multi-agent training can unlock unsupervised reasoning capabilities in RL, yielding performance gains that rival supervised methods without the need for costly annotations.
Subtracting information from the student rather than adding it to the teacher can yield performance improvements that rival those achieved with privileged data.
U-OPSD enables LLMs to self-improve without any external supervision, achieving up to 10.7% performance gains on challenging reasoning benchmarks.
VCSD achieves up to a 4.5% performance boost over traditional methods by distilling knowledge without requiring external teachers or visual cues.