Search papers, labs, and topics across Lattice.
This study quantitatively analyzes the implicit assumptions language models make about urban environments by evaluating anonymized profiles of real cities across 40 indicators and seven domains. By employing a robust methodology that includes constrained probability-based ratings and lineage-aware aggregation, the research reveals a consistent preference among models for urban profiles characterized by larger developed areas, rapid growth, and extensive infrastructure. The findings highlight how geographic differences in model outputs diminish when accounting for city scale and development, providing a clearer understanding of how language models conceptualize urbanity.
Language models exhibit a surprising bias towards cities with expansive infrastructure and rapid growth, revealing their implicit urban assumptions.
Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and function. We measure those assumptions without naming places. Ten open-weight checkpoints rate anonymized profiles derived from real morphological urban centres across 40 audited indicators and seven domains. The design combines constrained probability-based ratings, prespecified reliability screens, lineage-aware aggregation, multiple population weightings, an independent replication sample, and whole-profile validation. The clearest shared tendency favours urban profiles with larger developed area, faster recent growth, greater mapped infrastructure and non-residential capacity, and less sparse form. Most eligible directions recur in the replication data, and direct ratings of complete profiles show moderate agreement with the indicator-wise construction. Geographic differences shrink after accounting for city scale and development, while reliably measured paired tasks indicate that typicality and desirability are often closely aligned. The framework makes an otherwise vague notion of what models regard as an ordinary city empirically traceable. The resulting evidence delineates a shared yet model-dependent portrait of the city through the lens of language models.