Search papers, labs, and topics across Lattice.
This study evaluates the ability of Vision-Language Models (VLMs) to classify building typologies from Google Street View images, comparing their predictions to those made by human experts in civil engineering and architecture. By employing various scaling strategies and prompting techniques, the research reveals that Chain-of-Thought prompts enhance model performance stability, achieving an average accuracy of around 70%. Notably, the findings indicate that while VLMs rely heavily on visual indicators, human experts incorporate broader contextual knowledge, highlighting the complementary roles of AI and human reasoning in urban analysis.
VLMs can achieve 70% accuracy in building typology classification, but they often miss the broader context that human experts consider crucial.
This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Storeys from Google Street View (GSV) images. Predictions generated by VLMs are compared with inference by human experts (civil engineers and architects) as a source of manually labelled ground-truth data. We evaluate several state-of-the-art VLMs, including GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash. By applying different scaling strategies and prompting techniques, we found that Chain-of-Thought prompts provide an overall more stable model performance. We also investigate the reasoning behind VLMs'building-typology predictions by examining the probabilities of keywords appearing in AI explanations. This enabled us to analyse patterns in these reasonings and identify key themes driving both agreements and disagreements between VLM and expert labels. We find that AI tends to focus on visual indicators, whereas human experts place greater emphasis on broader contextual cues and domain knowledge, in addition to visual cues. Overall, VLM can approximate experts'capability in building-typology classification at scale, with an average accuracy of approximately 70%. The study demonstrates the VLM's potential for AI automation in tasks that require pattern recognition and object identification in an urban context. AI have the potential to serve as complementary and collaborative tools for urban analysis, leveraging their strengths in understanding visual patterns. This study contributes to the exploration of the efficiency and scalability of AI visual prediction and provides insights into the reasoning processes that could support automation processes in urban analysis and prediction.