Search papers, labs, and topics across Lattice.
SiZeUp introduces a novel method for creating large-scale 3D urban proxy models from calibrated oblique aerial imagery by employing a height-from-footprint representation that simplifies the reconstruction process. The approach utilizes an ordinal depth consistency loss to ensure the relative depth ordering aligns with predictions from a monocular depth model, enabling efficient height estimation without the need for dense point cloud reconstruction. This results in a significant speedup of 23-52 times over existing methods while preserving proxy-level coverage and volume consistency, making it highly effective for urban modeling applications.
Achieving a 23-52x speedup in 3D urban modeling, SiZeUp leverages ordinal depth consistency to transform aerial imagery into accurate proxies without the pitfalls of traditional depth estimation.
We present SiZeUp, a fast and scalable approach for constructing large-scale 3D urban proxy models directly from calibrated oblique aerial imagery. Our method adopts a height-from-footprint representation, reducing 3D building abstraction to a low-dimensional optimization problem in which building footprints are extruded by a single height parameter. To enable efficient and robust height estimation, we introduce an ordinal depth consistency loss that enforces agreement between the relative depth ordering of rendered proxies and depth priors predicted by a monocular depth model. This is realized through a differentiable renderer that maps parametric building proxies into multi-view depth images, allowing gradients to be propagated from depth supervision to building heights. Our ordinal formulation produces stable optimization in practice and avoids explicit feature matching or dense point cloud reconstruction. Rather than relying on metric depth, which can be unreliable under monocular scale ambiguity, our ordinal depth consistency loss operates on relative depths, providing a more reliable signal across views. Combined with an efficient dynamic view selection, our approach achieves a 23-52$\times$ speedup over state-of-the-art proxy reconstruction pipelines while maintaining comparable proxy-level coverage and volume consistency, making it well suited for large-scale urban modeling tasks.