Search papers, labs, and topics across Lattice.
This paper surveys the detection of AI-generated videos (AIGC-V), shifting the focus from traditional artifact-centric methods to a more nuanced approach called Factual Fidelity Verification, which assesses the alignment of video content with real-world facts. By introducing a Vision-Language Dual-View taxonomy, the authors categorize existing detection techniques into a structured framework that encompasses intrinsic cue analysis, spatiotemporal consistency modeling, cross-modal reasoning, and language-guided verification. The comprehensive review of 221 works not only highlights the evolution of AIGC-V detection but also outlines current challenges and future directions for achieving robust and trustworthy detection systems.
A shift from artifact matching to evidence-based semantic verification could redefine how we detect AI-generated videos.
The evolving realism of AI-generated Videos (AIGC-V) is rapidly rendering traditional artifact-centric detection insufficient, necessitating a paradigm shift from low-level inspection to high-level semantic verification. This paper presents a comprehensive survey of AIGC-V detection, reframing the task as Factual Fidelity Verification, which asks whether the events, entities, and physical processes depicted in a video are consistent with real-world facts. To systematize this rapidly evolving field, we propose a Vision-Language Dual-View taxonomy that organizes existing methods into a hierarchical, four-layer landscape, spanning intrinsic cue analysis, spatiotemporal consistency modeling, cross-modal consistency reasoning, and language-guided world-level reasoning. This dual-view framing highlights a fundamental transition from artifact matching in traditional deepfake detection to evidence-based semantic verification enabled by vision-language models and agentic reasoning pipelines. Based on a systematic review of 221 works, we synthesize AIGC-V generation paradigms, survey the landscape of detection methods, and review evaluation metrics and benchmarks in line with proposed views. Finally, we discuss current challenges and identify promising directions toward robust, explainable, and trustworthy detection.