Search papers, labs, and topics across Lattice.
This paper introduces GenRec, a novel multi-view flow matching model that distinctly separates the reconstruction of visible pixels from the generation of unobserved pixels in novel view synthesis. By leveraging an observation mask and a monocular depth estimator, GenRec optimizes both RGB and scene-coordinate maps while ensuring that generative signals remain uncontaminated by reconstruction errors. The results demonstrate that GenRec achieves superior reconstruction fidelity in observed areas and outperforms generative baselines in perceptual quality for unobserved regions across multiple datasets, including RealEstate10K and DL3DV-10K.
GenRec achieves unparalleled reconstruction fidelity while enhancing perceptual quality in novel view synthesis by intelligently separating reconstruction from generation.
Generative novel view synthesis from sparse input images is rarely all reconstruction or all generation: pixels visible in some source view have a unique correct value modulated only by view-dependent shading, while pixels in disocclusions or beyond the captured volume admit a distribution of plausible completions. Existing generative novel-view-synthesis methods conflate these regimes under a single uniform loss, blurring the line between geometric fidelity and creative hallucinations even when scene geometry is injected through warped point clouds or projected depth. We introduce GenRec, a multi-view flow matching model that builds the reconstruction--generation split directly into its architecture, supervision, and gradient flow. Guided by an observation mask derived from the source cameras and a monocular depth estimator, a flow matching backbone jointly denoises RGB and scene-coordinate maps across all target views, while a pixel-space refinement stage restores high-frequency detail on observed pixels; the same mask gates supervision so regression signals do not contaminate the generative prior. Across RealEstate10K, DL3DV-10K, and Mip-NeRF~360, in both single-view extrapolation and two-view interpolation, GenRec attains the best reconstruction fidelity in observed regions while also surpassing purely generative baselines on perceptual quality in unobserved ones, showing the effectiveness of our approach.