Search papers, labs, and topics across Lattice.
This paper introduces MMArt, a comprehensive multimodal dataset consisting of 74,234 WikiArt paintings annotated with four distinct perspectives鈥攏arrative, formal, emotional, and historical鈥攁longside a unified caption. The study reveals that existing art datasets are limited by their single-perspective focus, which fails to capture the complexity of art interpretation, and establishes through various analyses that each perspective contributes uniquely to tasks such as retrieval and reconstruction. Key findings indicate that while narrative descriptions excel in retrieval tasks, formal descriptions are crucial for preserving compositional style, highlighting the necessity of a multi-perspective approach in art understanding.
No single perspective is sufficient for art interpretation; MMArt reveals that each perspective uniquely enhances understanding and retrieval of visual art.
Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with formal analysis, grounded historical interpretation, or affective characterization. We argue this is not only a model but also a dataset limitation. Existing art datasets are single perspective resources, where no dataset provides narrative, formal, emotional, and historical perspectives simultaneously for the same artworks. We introduce MMArt, a large-scale dataset of 74,234 WikiArt paintings, each annotated with four independently annotated perspectives plus a harmonized unified caption, produced by specialized vision-language models or human annotation and validated through complementary quality evaluations. Two complementarity analyses establish that perspectives encode genuinely distinct information. A generative analysis shows that formal analysis descriptions best preserve compositional style, and historical descriptions carry strong affective signal in reconstructed images. A discriminative retrieval analysis reveals task-asymmetry: narrative descriptions drive retrieval (R@1 = 44.0%), while formal descriptions, strongest for reconstruction, are nearly nondiscriminative at retrieval scale (R@1 = 7.8%). Leave-one-out analysis further confirms that historical descriptions are the least replaceable perspective across both tasks. Together, the two analyses establish that no single perspective suffices for all tasks, directly motivating MMArt multi-perspective design. The dataset, code, and additional information are available at https://shuaiwang97.github.io/MMArt/.