Search papers, labs, and topics across Lattice.
3
29
6
6
Leading MLLMs falter on the new VideoGAIA benchmark, scoring under 60% accuracy in complex, multi-turn video understanding tasks.
MemDreamer narrows the performance gap with human experts in long video understanding to just 3.7 points while processing only 2% of the full context.
GPT-4o now has open-source competition: Ming-Omni matches its modality support in a single, unified model capable of perception and generation across image, text, audio, and video.