Search papers, labs, and topics across Lattice.
7
0
8
5
Simply adding more multimodal environments can hinder agent performance, but targeted diversity and structured difficulty can transform training outcomes.
Catastrophic collapses in tool-use performance can be mitigated by strategically interleaving supervised fine-tuning with reinforcement learning.
CoT reasoning boosts verbal reasoning but falters in visual tasks, revealing a critical gap in multimodal AI capabilities.
Agentic environments are not just backdrops for LLMs; they are pivotal in shaping agent evolution and capabilities.
Instance-level experiential knowledge boosts LLM tool use performance significantly more than abstract knowledge, reshaping how we approach knowledge integration in AI.
Squeezing the most out of your MLLM's visual budget is now possible: ResAdapt learns to allocate visual tokens intelligently *before* encoding, boosting performance by 15% while processing 16x more frames at the same cost.
MLLMs can now reason about streaming video with significantly improved accuracy and reduced output length thanks to a novel memory-anchored framework that overlaps watching and thinking.