Search papers, labs, and topics across Lattice.
AnimationBench is introduced as a new benchmark specifically designed to evaluate the performance of image-to-video generation models on animation-style content, addressing the limitations of existing benchmarks focused on realistic videos. It operationalizes the Twelve Basic Principles of Animation and IP Preservation into measurable evaluation dimensions, alongside broader quality dimensions like semantic consistency and motion rationality. Experiments demonstrate that AnimationBench aligns well with human judgment and reveals animation-specific quality differences that are missed by realism-oriented benchmarks.
Current video benchmarks fail to capture the nuances of animation, but AnimationBench fills the gap by rigorously evaluating character consistency, motion, and style.
Video generation has advanced rapidly, with recent methods producing increasingly convincing animated results. However, existing benchmarks-largely designed for realistic videos-struggle to evaluate animation-style generation with its stylized appearance, exaggerated motion, and character-centric consistency. Moreover, they also rely on fixed prompt sets and rigid pipelines, offering limited flexibility for open-domain content and custom evaluation needs. To address this gap, we introduce AnimationBench, the first systematic benchmark for evaluating animation image-to-video generation. AnimationBench operationalizes the Twelve Basic Principles of Animation and IP Preservation into measurable evaluation dimensions, together with Broader Quality Dimensions including semantic consistency, motion rationality, and camera motion consistency. The benchmark supports both a standardized close-set evaluation for reproducible comparison and a flexible open-set evaluation for diagnostic analysis, and leverages visual-language models for scalable assessment. Extensive experiments show that AnimationBench aligns well with human judgment and exposes animation-specific quality differences overlooked by realism-oriented benchmarks, leading to more informative and discriminative evaluation of state-of-the-art I2V models.