Search papers, labs, and topics across Lattice.
12
0
13
15
Rubric-based scoring can transform how vision-language models ground their responses, leading to substantial improvements in reasoning accuracy.
Kimi K3's innovative architecture achieves a 2.5x scaling efficiency improvement, enabling robust performance across diverse long-horizon tasks.
Atomic visual perception in MLLMs is largely unsolved, with no model surpassing 60% accuracy on a new benchmark designed to isolate perceptual capabilities.
Current video generation models struggle with law-grounded reasoning, with the best achieving only 47% on the new Apple-PI benchmark.
VideoChat3 achieves unprecedented generalization in video understanding while maintaining high efficiency, outperforming larger models with just 4 billion parameters.
ViQ achieves a groundbreaking balance between semantic richness and detail in visual representations, enabling efficient multimodal training without sacrificing quality.
Spatial reasoning can be transformed from isolated frame predictions to dynamic scene understanding, significantly boosting performance in multi-view and video tasks.
Ditching modular architectures unlocks surprisingly competitive vision-language performance, proving that end-to-end pixel-to-word models can rival traditional approaches at scale.
Current audio-visual generation models struggle to maintain coherence and alignment when scaling to minute-long content, a problem exposed by the new LongAV-Compass benchmark.
Leaderboard-topping video models are still surprisingly brittle, failing on basic video reasoning tasks unless given the right textual cues.
Forget dialogue summaries – FileGram builds user profiles directly from atomic file-system actions, unlocking a richer, more privacy-preserving approach to agent personalization.
Today's best multimodal models can only solve half of compositional visual tool-use tasks, revealing a critical gap in their ability to plan and execute complex, multi-step visual reasoning.