Search papers, labs, and topics across Lattice.
7
0
8
7
Visual updates in DeltaV cut token generation by over half while boosting reasoning accuracy, challenging the need for full-image outputs in multimodal models.
Achieving state-of-the-art performance in in-image machine translation, UniTranslator reveals a powerful synergy between translation understanding and image generation.
Treating negative samples differently based on their similarity to positives leads to a 13.1% boost in retrieval performance on complex queries.
Doc-V* demonstrates that an agentic approach to multi-page document VQA, using active navigation and structured memory, can significantly outperform retrieval-augmented generation, especially in out-of-domain scenarios.
OmniJigsaw reveals a "bi-modal shortcut phenomenon" in joint audio-visual integration, demonstrating that naive fusion can be surprisingly ineffective and highlighting the importance of carefully designed cross-modal training strategies.
Unleashing powerful reasoning in OLLMs doesn't require expensive training data or compute – just clever guidance from existing Large Reasoning Models.
By jointly training a keyframe sampler with an MLLM, MSJoE achieves state-of-the-art accuracy in long-form video understanding while significantly reducing computational cost.