Search papers, labs, and topics across Lattice.
This paper introduces RA-CLIPScore, a novel evaluation metric for generative models that enhances interpretability by measuring spatial distribution alignment and attribute-wise analysis. By employing dual prompts to decouple competing attributes and utilizing local patch tokens, RA-CLIPScore provides a more nuanced assessment of generated images compared to conventional metrics like FID and existing CLIP-based methods. Extensive experiments reveal that RA-CLIPScore not only offers robust evaluations under challenging conditions but also aligns more closely with human perception of visual diversity, highlighting spatial biases in generative models.
RA-CLIPScore reveals spatial biases in generative models, offering a more interpretable evaluation that aligns with human perception of visual diversity.
Generative models can produce images nearly indistinguishable from real data, yet rigorous and interpretable evaluation remains challenging. Conventional metrics such as FID provide only scalar scores with limited diagnostic insight. Widely adopted CLIP-based metrics enable semantic evaluation beyond simple training class labels, but inherit limitations from CLIP's training paradigm that restrict attribute-wise analysis. We propose RA-CLIPScore, a novel metric that mitigates these issues and extends CLIP-based evaluation to spatial distribution alignment, measuring whether generated objects adhere to the positional priors found in the training data. RA-CLIPScore introduces dual prompts to decouple competing attributes and leverages local patch tokens to capture fine-grained regional semantics. We evaluate image generative models on their ability to match both attribute and spatial distributions of the training data. Extensive experiments show that RA-CLIPScore provides more robust and interpretable evaluations than prior methods, particularly under distribution misalignment or partially irrelevant textual attributes. We further demonstrate how it reveals spatial biases in generative models. User evaluations confirm that Regional Single Attribute Divergence based on our RA-CLIPScore aligns more closely with human perception of visual diversity than existing semantic metrics.