Search papers, labs, and topics across Lattice.
This study investigates the phenomenon of temporal entanglement in museum datasets, where pretrained image representations may reflect institutional biases rather than true historical timelines. By framing artwork dating as an uncertainty-aware regression task using frozen image embeddings, the authors evaluate the performance of various pretrained vision models on a controlled corpus of artworks. The findings reveal that Vision-Language Models (VLMs) outperform visual self-supervised baselines in extracting temporal information, although this knowledge is influenced by underlying biases.
VLMs reveal hidden temporal insights in art history, but their biases could mislead interpretations of historical timelines. WHY_IT MATTERS: This research challenges the reliability of pretrained models in historical contexts, highlighting the need for critical evaluation of data representation in AI systems.
Museum and archival datasets do not mirror historical artistic production, but materialize the contingent histories of collecting, preservation, cataloging, and digitization. This has direct consequences for interpreting pretrained image representations: they may appear to encode historical time while actually encoding the institutional conditions under which objects become visible as data. We describe this phenomenon as temporal entanglement and investigate it by formulating artwork dating as an uncertainty-aware regression task over frozen image embeddings. We evaluate several pretrained vision models on a temporally controlled Wikidata corpus of artworks. Our results show that these models contain usable temporal information, with Vision-Language Models (VLMs) outperforming purely visual self-supervised baselines. However, a qualitative analysis indicates that this temporal knowledge is shaped by various biases.