Search papers, labs, and topics across Lattice.
This paper introduces FIRM-Video, a novel framework for text-to-video reward modeling that employs a checklist-driven approach to enhance evaluation accuracy while maintaining inference efficiency. By implementing a check-before-score principle, the framework verifies dimension-specific criteria against temporal visual evidence, leading to more reliable and interpretable reward assessments. The results demonstrate that the Qwen3-VL-based model achieves superior performance on the FIRM-Video-Bench, outperforming existing methods in key evaluation metrics across various video generators.
FIRM-Video reveals that a checklist-driven approach can significantly enhance the reliability of text-to-video reward modeling, achieving state-of-the-art performance in evaluation metrics.
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.