Search papers, labs, and topics across Lattice.
3
0
2
Achieving parallel region captioning with multimodal diffusion models could redefine efficiency benchmarks in visual perception tasks.
MLLMs can revolutionize video understanding by integrating watching, remembering, and reasoning into a cohesive framework that addresses long-range dependencies and sparse evidence.
LoomVideo achieves state-of-the-art video generation and editing efficiency with a compact architecture that accelerates inference speed by over 5 times compared to larger models.