Search papers, labs, and topics across Lattice.
This paper introduces GATO-Vid, a novel training-free and gradient-free method for spatially grounded text-to-video generation that addresses the limitations of existing gradient-based optimization techniques. By employing an analytical approach to cross-attention scores, GATO-Vid achieves precise spatial guidance without the computational burden of backward passes, making it suitable for large-scale architectures. Experimental results show that GATO-Vid significantly enhances localization accuracy compared to current baselines while maintaining low computational overhead.
GATO-Vid achieves superior spatial localization in text-to-video generation without the computational costs of traditional gradient-based methods.
Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant challenge. While existing training-free methods produce solid overall results in spatially grounded generation, \ie, placing a specific object in a designated location, they rely on gradient-based optimization techniques that incur prohibitive computational overhead, a bottleneck amplified in modern large-scale architectures. To address this limitation, we present Gradient-free Analytical Trajectory Optimization Video Generation (GATO-Vid), a novel training-free and gradient-free approach for precise spatial guidance. Rather than relying on costly backward passes, we introduce an alternative cross-attention score and solve it analytically to obtain an exact, closed-form solution. To use our analytical solution, we propose an on-the-fly injection mechanism tailored to the topological manifold of the transformer's latent space. Our experiments demonstrate that GATO-Vid significantly outperforms existing baselines in localization accuracy while introducing minimal computational overhead.