Search papers, labs, and topics across Lattice.
This paper addresses the instability of camera-controlled novel view synthesis during inference, particularly under significant camera motion and extended generation horizons. By decomposing camera motion into small autoregressive steps, the authors effectively limit geometric distortion and error accumulation, leading to improved stability. Their method, CamTrol++, demonstrates enhanced temporal and geometric consistency, as well as better downstream 3D reconstruction quality across multiple datasets, without the need for retraining the diffusion model.
Small autoregressive camera motions can drastically enhance the stability of novel view synthesis, preventing geometric distortion and error accumulation.
Training-free, camera-controlled novel view synthesis from a single image using pre-trained video diffusion models often becomes unstable under large camera motion and long generation horizons. Existing approaches commonly combine several inference-time components, making it unclear which design choices are most important for stability. We show that the main source of stability is simple. Decomposing camera motion into small autoregressive steps limits per-step geometric distortion and reduces error accumulation. A controlled camera-step study shows that performance remains stable for small motions and degrades more strongly as the per-step motion approaches $18$-$20^\circ$. We further evaluate geometry-constrained spatial attention and low-frequency appearance anchoring as supporting refinements, together with an efficient registration-free warping pipeline. Across RealEstate10K and MegaScene, CamTrol++ improves temporal and geometric consistency, downstream 3D reconstruction quality, and generation efficiency over training-free baselines. The method remains effective for 56-frame generation and under substantial controlled depth corruption. These results show that careful control of camera motion at inference time can substantially improve the stability of camera-controlled novel view synthesis without retraining or modifying the diffusion backbone.