Search papers, labs, and topics across Lattice.
This study investigates the relationship between visual representation richness and human alignment in urban engagement, utilizing 61 first-person city-walk videos segmented into over 50,000 clips across four modalities. While spatiotemporal video features initially show the strongest alignment with human engagement, temporally averaged images (TAIs) outperform video in binary classification tasks, challenging the assumption that richer representations are always superior. An independent study confirms that TAIs match human judgment accuracy for identifying engaging moments, particularly in composition-driven scenes, suggesting that simpler representations can be more effective in certain contexts.
Richer visual representations don't always lead to better human alignment; sometimes, simpler methods like temporally averaged images outperform full video in identifying engaging moments.
We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouTube segmented into over 50,000 ten-second clips and represented across four modalities: spatiotemporal video features, temporally averaged images (TAIs), audio embeddings, and text-based semantic descriptions. Spearman correlation analysis reveals the expected ordering along the temporal-richness continuum, with video features showing the strongest continuous alignment. However, this ordering breaks down under binary classification of high- versus low-engagement moments (the paradigm most commonly used to train perceptual scoring models), where TAIs consistently match or outperform video across most classifiers and quantile thresholds. An independent two-alternative forced-choice study on Amazon Mechanical Turk confirms that this parity reflects human judgment: participants identified engaging moments with comparable accuracy from TAIs and full video clips, while text performed substantially worse and audio remained near chance. Gap analysis reveals a functional dissociation: video features are advantaged in activity-driven scenes with dynamic content, whereas TAIs better align with human judgments in composition-driven scenes dominated by stable spatial structure. These findings challenge the assumption that richer representations are inherently more human-aligned, and suggest that perceptually grounded temporal compression can be a principled alternative to full video encoding.