Search papers, labs, and topics across Lattice.
To evaluate whether vision-language models (VLMs) can reverse-engineer generation-critical controls from synthetic media, the authors establish VI-Bench, a benchmark comprising 900 human-verified AIGC videos derived from 16.1 million real-world prompts across five control axes: subject, action, scene, style, and camera. This capability is essential for generative editing, asset reuse, and auditing prompt leakage vulnerabilities, yet it demands replay-stable control parameters rather than passive visual captions. Evaluating 18 frontier proprietary and open-source VLMs shows that even the best model achieves an Inversion Score of only 0.632, frequently generating semantically plausible prompts that catastrophically fail to regenerate the reference video's camera dynamics and multi-shot composition.
Even top-tier VLMs score just 0.632 when reverse-engineering AI video prompts, frequently hallucinating plausible descriptions that completely fail to reproduce the source video's camera and stylistic controls.
Recent advances in video generation have made prompt-based control increasingly central to AIGC video generation. Prompts specify what a video should depict and how it should be represented, controlling factors such as visual style or camera behavior. Understanding this recoverability is important both for creative reuse and editing, and for assessing prompt leakage risks. However, existing video understanding benchmarks do not measure this capability: a caption may describe what is visible, but a replayable prompt must recover the generation-relevant controls needed to reproduce the video. To address this gap, we introduce VI-Bench, a benchmark built from 16.1 million real-user prompts and 900 human-verified AIGC videos. VI-Bench spans three progressively harder settings, namely single-shot semantic grounding, control over style and camera behavior, and multi-shot compositional inversion, and evaluates five generation-critical dimensions: subject, action, scene, style, and camera. We evaluate 18 representative VLMs, including 2 proprietary and 16 open-source models on VI-Bench, using an Inversion Score that measures prompt-level alignment with the original prompt and video-level fidelity of the regenerated video. The results reveal substantial limitations: even the strongest model achieves only 0.632 on Inversion Score, performance degrades sharply as samples require richer control and multi-shot reasoning, and models often produce plausible prompts whose regenerated videos deviate from the reference. These findings show that video prompt inversion is a distinct and under-evaluated capability requiring models to transform visual understanding into replay-stable generative control.