Search papers, labs, and topics across Lattice.
12
0
8
7
Visual updates in DeltaV cut token generation by over half while boosting reasoning accuracy, challenging the need for full-image outputs in multimodal models.
Switch-Reasoner reveals that adaptive reasoning selection can enhance MLLM performance by reducing unnecessary cognitive load while maintaining accuracy.
UI-MOPD achieves a remarkable balance between retaining existing capabilities and adapting to new platforms, with task success rates that challenge conventional approaches in GUI agent learning.
Real-world GUI agents can achieve over 72% success in complex mobile tasks by leveraging a unique hybrid training approach that integrates real-device execution and adaptive learning from failures.
Achieving state-of-the-art performance in in-image machine translation, UniTranslator reveals a powerful synergy between translation understanding and image generation.
RS-Gen achieves state-of-the-art performance in image generation by autonomously identifying and resolving logical issues and knowledge gaps in real-time.
Treating negative samples differently based on their similarity to positives leads to a 13.1% boost in retrieval performance on complex queries.
STAR Allocation enables targeted policy updates in text-to-image generation, leading to unprecedented improvements in semantic alignment and text rendering.
Doc-V* demonstrates that an agentic approach to multi-page document VQA, using active navigation and structured memory, can significantly outperform retrieval-augmented generation, especially in out-of-domain scenarios.
VideoLLMs can now watch and think *simultaneously*, achieving 15x faster response times and improved accuracy on video understanding tasks.
Unleashing powerful reasoning in OLLMs doesn't require expensive training data or compute – just clever guidance from existing Large Reasoning Models.
By jointly training a keyframe sampler with an MLLM, MSJoE achieves state-of-the-art accuracy in long-form video understanding while significantly reducing computational cost.