Search papers, labs, and topics across Lattice.
This paper introduces PRISM, a novel approach that enables video generators to directly discriminate preferences from noisy latents without relying on pixel-based reward models. By employing a lightweight Query-based Aggregation head on a frozen video diffusion backbone, PRISM achieves state-of-the-art preference accuracy and demonstrates significant noise-robustness, facilitating early-stage Best-of-$N$ sampling. The findings reveal a strong correlation between the generative performance of the backbone and its evaluative capabilities, suggesting a pathway for self-improving video generation models.
Unlocking preference representation from noisy latents, PRISM achieves state-of-the-art accuracy while drastically reducing computational costs in video generation.
Evaluating video generation with clean, pixel-based reward models disconnects evaluation from the noisy diffusion process and incurs massive VAE decoding costs. In this paper, we challenge this paradigm by asking a fundamental question: Can a powerful video generator inherently discriminate preferences directly from noisy latents? To answer this, we introduce \textbf{PRISM} (\textbf{P}reference \textbf{R}epresentation in \textbf{I}ntermediate \textbf{S}tates of Diffusion \textbf{M}odels). PRISM employs a lightweight Query-based Aggregation head with a frozen video diffusion backbone to decode preference signals from noisy latents. Surprisingly, PRISM not only achieves SOTA preference accuracy but also unlocks strong noise-robustness, which enables early-stage Best-of-$N$ sampling. This allows for filtering suboptimal candidates at the very beginning of denoising, drastically reducing computation while boosting video quality. We also reveal a strong positive correlation between a backbone's generative performance and its inherent evaluative power, enabling self-improving video backbones.