Search papers, labs, and topics across Lattice.
6
0
8
7
Even the best vision-language models struggle with reliable evaluation of computer-using agents, but OS-Shepherd models offer a low-cost solution that matches their performance.
SoftSkill achieves up to 42.1 points improvement on LiveMath by transforming lengthy Markdown skills into a few powerful virtual tokens.
Training with large block sizes cripples reasoning performance, but a novel curriculum approach unlocks strong reasoning capabilities in diffusion models.
Social intelligence may require more than just reasoning power: a 7B model trained with SAVOIR beats GPT-4o and Claude-3.5-Sonnet on social interaction tasks.
STRATAGEM reveals that selectively reinforcing reasoning trajectories can dramatically enhance a model's ability to transfer reasoning skills across diverse tasks, especially in complex mathematical scenarios.
LLMs struggle to automatically apply learned procedures or avoid failed actions without explicit reminders, achieving only up to 66% on a new implicit memory benchmark.