Search papers, labs, and topics across Lattice.
This study explores how AI models autonomously categorize artistic works across multiple modalities鈥攖ext, audio, image, and video鈥攗sing a self-supervised framework that projects these modalities into a shared 256-dimensional embedding space. By applying iterative clustering, the researchers reveal significant divergences between AI-generated aesthetic clusters and human emotional categorizations, highlighting the unique ways AI interprets artistic media. The findings underscore the potential for AI to enhance media organization and retrieval processes, particularly in applications like Retrieval-Augmented Generation (RAG) and automated data labeling.
AI's aesthetic categorization of art diverges significantly from human emotional responses, revealing a unique interpretative framework that could transform media organization.
Aesthetics are an important part of the symbolism of artistic works. Although subjective, humans categorize art based on the emotion evoked regardless of modality. What remains under-explored is how AI models form their own aesthetic categorization of human-produced media without explicit labels or cross-modal supervision. We present a self-supervised framework that projects four modalities (text, audio, image and video) into a shared 256-dimensional embedding space and applies iterative clustering to discover aesthetic structure. We discuss the divergence between AI-generated cluster assignments and human affective register labels on a weakly supervised multimodal dataset. This work has applications in understanding how AI structures cross-modal similarity, organizing heterogeneous media collections for Retrieval-Augmented Generation (RAG), and automated data labeling.