Search papers, labs, and topics across Lattice.
This paper introduces AesCanvas, a comprehensive dataset and benchmark designed to evaluate both aesthetic critique and contextual suitability of images across various domains. By providing two components鈥擟ritiqueCanvas for multi-dimensional critique and ContextCanvas for assessing contextual appropriateness鈥攖he authors highlight the limitations of existing benchmarks that focus solely on intrinsic visual quality. The results indicate a significant divergence between the capabilities of models in generating critiques versus making context-sensitive judgments, underscoring the need for a more nuanced approach to aesthetic evaluation in MLLMs.
AesCanvas reveals that aesthetic specialization does not guarantee contextual suitability, challenging the assumptions about model performance in image assessment.
Recent advances in Multimodal Large Language Models (MLLMs) have extended Image Aesthetic Assessment (IAA) beyond scalar scores toward interpretable critique and guidance. Yet existing benchmarks mainly assess intrinsic visual quality or fixed domain criteria, leaving open whether an appealing image is appropriate for a specific purpose, audience, cultural setting, or domain convention. We introduce AesCanvas, a unified suite with two complementary components: CritiqueCanvas with 519,136 instruction-response pairs from 54,300 images supports long-form, multi-dimensional critique across photography, painting, and virtual imagery, whereas ContextCanvas with 301 expert-reviewed use scenarios evaluates contextual aesthetic suitability in realistic use scenarios. Under a unified protocol, we evaluate closed-source frontier, open-weight general, and aesthetic-specific MLLMs. Results reveal a clear separation between critique generation and context-sensitive judgment: reference-based lexical and semantic metrics only partially capture critique quality, while aesthetic specialists remain competitive on selected critique metrics yet substantially lag strong general-purpose MLLMs on ContextCanvas. Further analyses show that aesthetic specialization does not reliably transfer to contextual suitability and that model decisions may fail to track or ground themselves in decisive contextual visual cues. These findings establish culturally situated, evidence-grounded suitability as a distinct objective for aesthetic modeling.