Search papers, labs, and topics across Lattice.
This paper introduces the Self-Generative-Understanding (SGU) framework, which evaluates unified multimodal models (UMMs) by integrating their generative and discriminative capabilities in a cohesive manner. By employing a semantic closed-loop challenge that involves perceiving an image, generating a textual description, reconstructing visual context, and reasoning over the output, SGU provides a zero-cost evaluation method that does not require new annotations. The findings reveal that high-performing UMMs often struggle with reasoning over their own generated contexts, highlighting critical limitations that traditional separate evaluations fail to capture.
Unified multimodal models may excel in generation and understanding, but they often falter when reasoning about their own outputs, revealing hidden weaknesses in their capabilities.
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs). In this work, we propose Self-Generative-Understanding (SGU), a novel, annotation-free evaluation framework that probes the integrated capabilities of unified models through a semantic closed-loop challenge. Without requiring new annotations, SGU leverages the dual understanding-and-generation abilities of UMMs by asking them to first perceive an image and produce a textual description, subsequently reconstruct a visual context based on that description, and finally perform reasoning over the self-generated output. This pipeline provides a zero-cost testbed that yields an integrated performance score specifically tailored for evaluating UMMs as unified systems. Extensive experiments show that even high-performing UMMs often struggle to reason over their own generated contexts, revealing limitations that are not captured by separate evaluations of understanding or generation alone. Our work provides a complementary holistic evaluation framework and offers a foundation for benchmarking the development of next-generation unified multimodal models.