Search papers, labs, and topics across Lattice.
This paper introduces the "Agentic Self-Improvement" framework to enhance control and reliability in image-to-video (I2V) models by transforming video synthesis into a goal-directed optimization process. The approach employs a two-stage method where a multimodal Large Language Model refines prompts through automated evaluations, followed by Bayesian optimization to co-optimize stochastic seeds and CFG scales. The results show a significant improvement in user preference for videos generated with this framework, achieving win rates of up to 69% over traditional unguided methods, thereby advancing the practicality of video generation technologies.
Videos generated through a novel agentic optimization framework are preferred by users over traditional methods, achieving a striking 69% win rate in preference studies.
Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperparameters to yield drastically different outputs often necessitating inefficient, brute-force trial-and-error processes. To address these limitations, we introduce the ``Agentic Self-Improvement"framework, which reframes video synthesis into a closed-loop, goal-directed optimization. Our framework systematically navigates the generation parameter space using a novel two-stage approach. In the first stage, an iterative prompt optimization loop uses a multimodal Large Language Model (mLLM) to refine the input prompt. This refinement implements two automated evaluations: Davidsonian Scene Graph (DSG) queries ensure semantic adherence, and Common Mistake Questions (CMQ) for artifact detection. At the second stage, we use Bayesian optimization to efficiently co-optimize stochastic seeds and CFG scales. This search is guided by a suite of quality metrics, including the novel Video-Text Adherence (VTA) score derived from the DSG and CMQ evaluations. Our framework significantly outperforms unguided search methods: in human preference studies, videos generated via our agentic approach were strongly preferred over baseline outputs, achieving win rates up to 69\%. This work provides a practical and extensible methodology for enhancing the predictability and control of state-of-the-art video generation models, moving the field beyond speculative curiosities toward reliable, production-ready tools.