Search papers, labs, and topics across Lattice.
4
0
6
8
VGIF-Score reveals that current video generation models struggle with complex instructions, providing a diagnostic lens to pinpoint where they succeed or fail.
Achieving lossless processing of 256K contexts, Keye-VL-2.0 transforms how we approach long-video understanding and agentic intelligence.
Current MLLMs can't find the lies hidden in their long image captions, struggling to pinpoint specific hallucinated words within detailed narratives.
You can now get state-of-the-art hepatocellular carcinoma diagnosis and captioning from whole slide images using a new MLLM with a topology-aware attention mechanism.