Search papers, labs, and topics across Lattice.
6
0
8
2
Light-Omni achieves a remarkable 12.1× speedup in video understanding while enhancing accuracy, redefining efficiency in agentic video processing.
EvoEmbedding outperforms leading embedding models by adapting its representations in real-time, revolutionizing long-context retrieval and agentic memory integration.
Fine-tuning on the new OmniVideo-100K dataset boosts model performance by over 20% in audio-visual reasoning tasks, revealing the power of structured scripts in enhancing multimodal understanding.
Current audio-language models are surprisingly bad at controlling and interpreting subtle vocal cues, failing in nearly half of situational dialogue scenarios.
Current MLLM benchmarks are missing the forest for the trees: Agentic-MME reveals that strong final-answer accuracy masks surprisingly poor tool use and planning in complex multimodal tasks.
Forget static, single-turn personalization – PersonaVLM unlocks long-term, evolving user alignment in MLLMs, even surpassing GPT-4o.