Search papers, labs, and topics across Lattice.
The authors present StreetDiff, a multi-view diffusion framework paired with the new Street360 dataset to tackle geometric inconsistencies and object duplication in complex urban scene generation. By decoupling global layout reasoning from local detail synthesis and injecting spherical-projection attention constraints into the denoising process, the approach enforces rigorous cross-view alignment without retraining the underlying diffusion backbone. Evaluations demonstrate that this explicit modeling of spherical correspondences substantially improves structural coherence and visual fidelity over prior multi-view diffusion baselines under camera rotation.
Multi-view diffusion models routinely hallucinate duplicate objects and warp geometry under camera rotation; enforcing spherical-projection attention constraints directly during denoising solves this across complex streetscapes without modifying backbone weights.
Multi-view diffusion models have shown strong performance in scenes with strong geometric priors and sparse semantics, such as indoor rooms or simple outdoor environments (e.g., fields, courtyards). However, they often fail to maintain cross-view consistency under camera rotation, especially in structurally complex urban environments. Without explicit modeling of spherical correspondence across views, existing approaches tend to produce object duplication, structural distortion, and layout inconsistency. To address this limitation, we propose StreetDiff, a multi-view diffusion framework that explicitly enforces cross-view alignment during denoising. StreetDiff introduces a Panorama--Perspective Synergy design to decouple global layout reasoning from local detail synthesis, and incorporates a Panorama Alignment Module (PAM) that establishes spherical-projection-based attention constraints across views. By injecting structured alignment constraints without modifying the diffusion backbone, our framework achieves robust cross-view coherence in challenging urban street scene generation tasks. In addition, we construct Street360, a large-scale HDR multi-view urban panorama dataset. Extensive experiments demonstrate that StreetDiff significantly improves structural consistency and visual fidelity compared to prior multi-view diffusion generation methods.