Search papers, labs, and topics across Lattice.
This paper introduces CultureVidBench, a novel benchmark designed to evaluate the cultural understanding of text-to-video (T2V) generation models, addressing a significant gap in existing assessments that focus primarily on perceptual quality and alignment. By curating 1,000 prompts across diverse cultural contexts, the benchmark emphasizes the importance of accurately representing culturally specific objects, actions, and rituals in generated videos. Evaluation of seven T2V models reveals that while they excel in semantic adherence and visual quality, they struggle to faithfully depict nuanced cultural details, especially from underrepresented regions.
Current T2V models may look good, but they often miss the mark on capturing the rich cultural nuances that define diverse contexts.
Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, and text-video alignment, but do not directly assess whether generated videos capture culturally specific objects, actions, rituals, visible text, or audio cues. We introduce CultureVidBench, a comprehensive benchmark for evaluating cultural understanding in T2V generation. CultureVidBench contains 1,000 curated prompts covering 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects organized into three categories: material culture, social practice&performance, and ritual&ceremony. Designed specifically for video generation, CultureVidBench emphasizes dynamic and multimodal cultural representation, including social interactions, ritual procedure, and culturally appropriate visible text and audio. We evaluate seven representative T2V models through human user studies and MLLM-based automatic assessment across cultural faithfulness, multimodal cultural rendering, semantic adherence, and perceptual quality. Results show that although current models achieve strong semantic adherence and visual quality, they often fail to faithfully capture fine-grained cultural details, particularly for underrepresented regions, rituals, and multimodal cultural cues.