Search papers, labs, and topics across Lattice.
This study investigates two approaches to learning visual representations for analyzing composition in artworks and photographs: a human-inspired method based on perceptual grouping and fine-tuned foundation models utilizing large-scale datasets. The findings reveal that while the human-inspired approach offers competitive performance and interpretability with frozen encoders, fine-tuning large self-supervised models leads to superior results but sacrifices interpretability and cross-domain generalization. This research highlights the trade-offs between interpretability and performance in compositional analysis, emphasizing the need for a balanced approach in future model development.
Fine-tuned foundation models can significantly outperform human-inspired methods in compositional analysis, but at the expense of interpretability and generalization.
Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it remains among the least formalized dimensions of visual understanding. Prior work highlights a persistent gap in learning meaningful compositional representations, attributing it to semantic bias and suggesting that human-inspired approaches may be key. We compare two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets. The human-inspired approach uses object-centric models for region-level decomposition and a graph attention network to capture spatial relationships between elements. Both paradigms are evaluated on composition score/category prediction, compositional image retrieval, and visual saliency detection. With frozen encoders, the human-inspired method achieves competitive performance while remaining interpretable. When sufficient data enables fine-tuning, large self-supervised models outperform significantly, but at the cost of interpretability and cross-domain generalization.