Search papers, labs, and topics across Lattice.
The AVA-Encoder introduces a novel auto-encoding framework that transforms videos into a structured Film Knowledge Graph (KG) representation, enabling agents to learn from high-quality human films. This structured representation captures essential entities, events, and their multimodal relationships, facilitating agentic reasoning and manipulation. Experimental results demonstrate a significant performance improvement, with AVA-Encoder achieving a 20.7-percentage-point gain over existing baselines, while also reducing the need for system-prompt tokens in policy training.
AVA-Encoder achieves a 73.1% relative improvement in video representation learning, enabling agents to produce cinematic-grade videos with far fewer resources.
Video creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a novel auto-encoding framework driven by agentic self-evolution to learn agent-native video representations. AVA-Encoder transforms a video into a Film Knowledge Graph (KG) representation and then reconstructs it back into video. This Film KG representation explicitly captures entities, events, assets, and their multimodal relationships in a structured form that can be easily understood, queried, and manipulated by agents. The reconstruction residual drives a dual-loop textual-gradient optimization framework that jointly improves the Film KG representation and the Agentic Video Encoder. Extensive experiments show that AVA-Encoder achieves a 20.7-percentage-point absolute gain, or a 73.1% relative improvement, over the strongest external baseline. In the controlled policy-only setting, its pseudo-trained Agentic Video Encoder policy also outperforms a carefully human-tuned policy while using 74.3% fewer shot-level and 70.1% fewer keyframe-level system-prompt tokens. We release the complete AVA-Encoder framework, a reliable agentic video reconstruction benchmark, and the first dataset of high-quality Film KG representations.