Search papers, labs, and topics across Lattice.
Affiliation:
4
0
5
This work found that certain attention blocks naturally form an Intrinsic Spatial Grounding Map that precisely locates reference subjects, and proposes Dual-phase Intrinsic Attention Leveraging (DIAL), a framework that uses these internal signals for both training and inference.
Achieving parallel region captioning with multimodal diffusion models could redefine efficiency benchmarks in visual perception tasks.
MLLMs can revolutionize video understanding by integrating watching, remembering, and reasoning into a cohesive framework that addresses long-range dependencies and sparse evidence.
LoomVideo achieves state-of-the-art video generation and editing efficiency with a compact architecture that accelerates inference speed by over 5 times compared to larger models.