Search papers, labs, and topics across Lattice.
9
0
12
8
Making the teacher's privileged context learnable end-to-end enables agents to evolve more efficiently, outperforming traditional methods with less than 30% of their rollout budget.
CED reveals that VLMs can be trained to prioritize evidence-based reasoning over language shortcuts, leading to more reliable visual understanding.
RefCaptioner not only outperforms existing models in video captioning but also enables precise grounding of visual elements to multiple reference images, enhancing factual accuracy.
Existing models mismanage tool use, but Beacon achieves a balance that enhances performance on complex tasks while preserving accuracy on simpler ones.
Current video generation models face a critical trade-off between faithfully executing keyframes and producing natural-looking videos, with performance degrading under increased keyframe density.
LaMem-VLA seamlessly integrates historical experience into VLA reasoning, enabling robots to perform complex tasks with improved contextual awareness.
DOPD reveals that intelligently routing supervision based on advantage gaps can significantly enhance capability transfer in distillation, outperforming conventional methods.
LLM multi-agent systems can achieve significantly higher accuracy at a fraction of the cost by learning to selectively delegate tasks instead of relying on rigid orchestration.
Language models are increasingly doing their real work in the "invisible" latent space, not the tokens we see.