Search papers, labs, and topics across Lattice.
This paper introduces D3-Omni, a novel benchmark designed to evaluate the performance of multimodal understanding models鈥攖ermed "OmniJudges"鈥攁cross text-to-image, text-to-video, and text-to-speech tasks. By employing a balanced and decoupled approach, D3-Omni addresses the limitations of existing benchmarks that often conflate distinct failure modes and overemphasize positive examples, revealing that even high-performing models struggle with modality-specific dimensions. The findings indicate that aggregate accuracy can mask significant capability gaps, underscoring the importance of a more nuanced evaluation framework for multimodal systems.
Even top-performing multimodal judges fail to reliably detect errors in specific modalities, revealing hidden blind spots in their evaluations.
Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as"OmniJudges"for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to overemphasize positive examples and to conflate distinct failure modes, so a judge may score well without recognizing failures while its capability gaps stay hidden. Motivated by this, we introduce D3-Omni, a balanced and decoupled benchmark for diagnosing fine-grained multimodal understanding, covering 53 orthogonal binary dimensions (17/22/14) and 10,671 samples (3,526/1,998/5,147) across the three tasks. Rather than re-generating outputs, which may leak information across dimensions, we fix verified fully positive seeds and derive negatives through controlled prompt rewriting and atomic, dimension-isolating perturbations. The resulting D3 design is Dual-balanced, which helps alleviate negative-sample scarcity and per-dimension label imbalance; Decoupled, so that each error is attributable to a single capability; and Dynamic, steering construction toward under-represented regions of the label distribution as generative models improve.The suite reaches near 1:1 per-dimension parity and a uniform distribution over all total-score levels. Under this balanced view, even strong OmniJudges tend to struggle on modality-related dimensions, to confirm satisfied requirements far more reliably than they detect violated ones, and to treat nominally distinct attributes as largely a single decision, suggesting that aggregate accuracy may hide systematic blind spots that a balanced and decoupled lens can help expose and, in turn, address.