Search papers, labs, and topics across Lattice.
This paper introduces CamWorldQA, a novel benchmark specifically designed for assessing the perceptual quality of camera-controlled world video generation, addressing the limitations of existing video quality assessment methods that focus on natural videos. The benchmark includes 720 generated videos from six generation methods, annotated with human-rated quality scores, highlighting unique perceptual characteristics such as viewpoint consistency and motion coherence. The proposed CWQA network, which integrates spatial, temporal, and optical flow features, outperforms traditional quality assessment approaches, showcasing its effectiveness in this new domain.
Existing video quality metrics fall short for camera-controlled generation, but CWQA sets a new standard by accurately predicting perceptual quality with a tailored approach.
Recent advances in generative video models have enabled camera-controlled world video generation, allowing models to synthesize videos under user-defined camera trajectories. However, existing video quality assessment (VQA) methods are mainly developed for natural videos and fail to capture the unique perceptual characteristics of camera-controlled generation, such as viewpoint consistency, motion coherence, and content preservation. In this work, we introduce CamWorldQA, the first benchmark for perceptual quality assessment of camera-controlled world video generation. CamWorldQA contains 720 generated videos produced by 6 representative generation methods from 20 diverse source videos under 6 camera trajectories, where each video is annotated with a human-rated perceptual quality score through subjective experiments. Furthermore, we propose CWQA, a no-reference quality assessment network with three complementary branches that extract spatial features, temporal motion features and optical flow features to jointly predict quality scores. Extensive experiments demonstrate that CWQA achieves superior performance over existing quality assessment methods on the CamWorldQA dataset.