Search papers, labs, and topics across Lattice.
7
0
10
6
Cooperative multi-agent training can unlock unsupervised reasoning capabilities in RL, yielding performance gains that rival supervised methods without the need for costly annotations.
Reallocating optimization effort based on reward saturation can boost performance by up to 9.2% in complex reasoning tasks.
Subtracting information from the student rather than adding it to the teacher can yield performance improvements that rival those achieved with privileged data.
Skills stabilize agent execution by transforming noisy trajectories into procedural anchors, but they can fail under brittle assumptions and incompatible contexts.
U-OPSD enables LLMs to self-improve without any external supervision, achieving up to 10.7% performance gains on challenging reasoning benchmarks.
VCSD achieves up to a 4.5% performance boost over traditional methods by distilling knowledge without requiring external teachers or visual cues.
Dense prediction can achieve state-of-the-art performance by leveraging generative pretraining without the overhead of RGB content generation.