Search papers, labs, and topics across Lattice.
PlaceSeek is a human-centered framework for urban outdoor place retrieval that enhances traditional geospatial methods by incorporating both functional and affective user intents. The system employs a Semantic Grounding Module to ensure that retrieved street-view images contain the necessary physical evidence for the intended activities, while an Affective Alignment Module re-ranks these candidates based on human perceptions of urban spaces. Evaluated on a dataset of 31,956 street-view locations in Milan, PlaceSeek achieved an impressive 88.0% Precision@5 and outperformed existing models like CLIP and SigLIP, demonstrating the importance of integrating both visual evidence and human affect in geospatial retrieval.
Users can now find urban outdoor places that resonate with their activities and emotions, not just their categories, thanks to a novel intent-aware retrieval framework.
People search for urban outdoor places not only by category or function, but also by what activities a place can support and how it is perceived. Existing geospatial retrieval remains largely POIcentric and metadata-driven, making it difficult to satisfy openended, affective, or activity-oriented needs. We present PlaceSeek, a human-centered outdoor place retrieval framework that maps natural-language queries to geolocated street-view imagery. PlaceSeek introduces an intent-aware retrieval mechanism that decomposes user queries into functional and affective sub-intents. A Semantic Grounding Module verifies whether candidate street-view results contain the physical evidence needed to support the intended activity, while an Affective Alignment Module re-ranks physically valid candidates using a LoRA-adapted vision-language model trained on human urban perception judgments. We evaluate PlaceSeek on 31,956 street-view locations in Milan across 10 naturallanguage queries annotated by five human evaluators. PlaceSeek achieves 88.0% Precision@5, a mean match score of 3.39/4.0, and 0.920 nDCG@5, outperforming CLIP, fine-tuned CLIP, SigLIP, and a VQA-based baseline. Ablation results show that physical grounding is essential for retrieval validity, while affective alignment improves ranking quality among physically valid candidates. These findings highlight that complex urban spatial queries require modeling both verifiable visual evidence and human perceptual preferences. PlaceSeek provides a potential framework for human-centered nextgeneration geospatial retrieval systems.