Search papers, labs, and topics across Lattice.
Xiaohongshu Inc.
8
1
10
4
StreamMind's innovative architecture not only enhances long-horizon video understanding but also slashes query-to-answer latency, setting a new standard for multimodal agent performance.
Harness-R1 achieves a remarkable 9.3 percentage point increase in task success rates by learning to edit agent harnesses in response to failure trajectories.
Synthesized egocentric videos can elevate real-robot task success rates by over 10% through enhanced training data.
LLM agents struggle to maintain performance in multi-day collaborative tasks, dropping significantly after just one environmental update, revealing a critical gap in adaptation to evolving real-world conditions.
Forget trajectory-level rollouts: MuSEAgent learns faster and reasons better by distilling past interactions into reusable, state-aware decision experiences.
Achieve significantly more accurate text and formula rendering with a training-free agentic workflow that injects glyph templates into latent spaces and attention maps of text-to-image models.
Current multimodal models are stuck in bi-modal interactions, but OmniGAIA and OmniAtlas offer a path towards truly omni-modal AI assistants capable of reasoning and tool use across video, audio, and images.
LLMs can learn to generate high-quality symbolic world models by interacting with a multi-agent system that provides adaptive, behavior-aware feedback, closing the gap between static validation and interactive execution.