Search papers, labs, and topics across Lattice.
16
0
10
10
A single unified model can outperform specialized systems across various computer vision tasks, all without the need for custom architectures.
Spatial reasoning can be transformed from isolated frame predictions to dynamic scene understanding, significantly boosting performance in multi-view and video tasks.
Spectral Forcing reveals that a simple low-pass filter can dramatically enhance diffusion model performance by effectively isolating signal from noise.
PermaVid achieves unprecedented long-term consistency in video generation, even after significant edits, by intelligently disentangling appearance and geometry in memory.
Prisma-World achieves unprecedented cross-view consistency in multi-agent video generation by leveraging a joint geometry-aware denoising process.
U4D reveals that leveraging spatial uncertainty can drastically enhance the quality of LiDAR scene synthesis, achieving unprecedented fidelity and coherence.
Ditching modular architectures unlocks surprisingly competitive vision-language performance, proving that end-to-end pixel-to-word models can rival traditional approaches at scale.
Spatial foundation models aren't as "all-round" as we thought: SpatialBench reveals surprising generalization gaps and the critical importance of domain alignment over naive data scaling.
LLaVA-OV-2's codec-stream tokenization lets it crush existing video-language models, especially in tasks requiring fine-grained temporal understanding of high-frequency motion.
A single framework now generates simulation-ready 3D assets for rigid, deformable, and articulated objects, unlocking new possibilities for embodied AI and physics-based simulation.
Zero-shot synthesis of articulated human-object interactions is now possible by treating diffusion-generated videos as supervision for 4D scene reconstruction, unlocking physically grounded interactions beyond rigid manipulation.
Unified multimodal models often *hurt* performance on multimodal understanding tasks, except for spatial reasoning, visual illusions, and multi-round reasoning, challenging the assumption that generation universally improves understanding.
Achieve 10% higher success rates in robotic manipulation tasks while speeding up inference by 1.5-1.8x by intelligently pruning visual tokens in multi-view Vision-Language-Action models.
A 1000x larger video reasoning dataset reveals early signs of emergent generalization, offering a new foundation for training and evaluating spatiotemporal AI.
Achieve SOTA joint audio-video generation with JavisDiT++ using just 1M public training examples, rivaling performance of models trained on proprietary datasets.
Turns out, skipping the boring parts of a video (like static backgrounds) makes your vision AI both faster and smarter, beating state-of-the-art models with less data.