Search papers, labs, and topics across Lattice.
Affiliation:, University of Washington
4
0
8
36
Relying on a single simulator in multi-agent RL leads to dangerous mode collapse, but innovative solutions can boost generalization and performance by up to 14%.
ZPPO reveals that embedding teacher responses in prompts rather than gradients can dramatically boost the performance of small student models on challenging tasks.
Fine-tuning on the new ProCUA-SFT dataset boosts UI-TARS 7B's performance from a dismal 8-10% to an impressive 45.0% on OSWorld tasks, highlighting the critical role of high-quality training data.
Multimodal models can now achieve state-of-the-art performance in real-world tasks like document understanding and audio-video comprehension with significantly reduced inference latency thanks to novel token-reduction techniques.