Search papers, labs, and topics across Lattice.
This paper introduces StreamAV-Bench, a novel benchmark specifically designed for evaluating streaming audio-video generation, addressing the limitations of existing benchmarks that focus on completed sequences. The framework includes two tracks: a progressive track for assessing instruction adherence and stability, and an interactive track for evaluating responsiveness and state management. Analysis of 13 representative generative systems reveals significant issues with temporal drift and responsiveness, highlighting critical areas for improvement in future models.
Current generative models struggle with temporal drift and responsiveness in streaming audio-video generation, as revealed by the new StreamAV-Bench benchmark.
Recent advancements in generative models are pushing video generation toward unbounded streaming audio-video generation for real-time interactive worlds. However, existing benchmarks primarily evaluate completed sequences and struggle to capture streaming properties. To bridge this gap, we introduce StreamAV-Bench, the first comprehensive benchmark tailored for streaming audio-video generation. StreamAV-Bench establishes a unified evaluation framework, including the progressive track for instruction adherence and long-horizon stability, and the interactive track for interactive response and state retention and reuse. With expert-verified evaluation cases across 32 fine-grained dimensions, we conduct an extensive evaluation of 13 representative systems. Our analysis reveals that current models suffer from temporal drift in progressive generation and responsiveness bottlenecks during interactive control. Based on a comprehensive failure analysis, we share insights to advance the development of native joint audio-video streaming models.