Search papers, labs, and topics across Lattice.
Affiliation:
3
0
7
This work found that certain attention blocks naturally form an Intrinsic Spatial Grounding Map that precisely locates reference subjects, and proposes Dual-phase Intrinsic Attention Leveraging (DIAL), a framework that uses these internal signals for both training and inference.
LoomVideo achieves state-of-the-art video generation and editing efficiency with a compact architecture that accelerates inference speed by over 5 times compared to larger models.
Achieve up to 4.87x faster diffusion sampling without retraining, and sometimes even *better* image quality, by intelligently planning the optimal denoising trajectory.