Search papers, labs, and topics across Lattice.
SceneReGen introduces a novel framework for single-image 3D scene reconstruction that bridges the gap between object-level generation and scene assembly by leveraging selective pose factorization. This method encodes each object's orientation directly into the generated mesh while estimating translation and scale from scene evidence, allowing for coherent placement of objects in a shared frame. On the 3D-FUTURE evaluation subset, SceneReGen outperforms existing methods in scene-level metrics and demonstrates strong potential for applications in autonomous driving and embodied AI scenarios.
SceneReGen achieves state-of-the-art performance in 3D scene reconstruction by seamlessly integrating object generation with scene assembly, redefining how we approach single-image reconstruction tasks.
Single-image 3D scene reconstruction must complete partially observed objects and place them coherently in a shared observation-aligned scene frame. Object-level generative priors offer strong completion ability, but their centered, scale-normalized outputs are typically expressed in an object frame, creating a fundamental representation gap between object generation and scene reconstruction. We introduce SceneReGen, a generative reconstruction framework that reinterprets scene reconstruction as the generation and assembly of complete object assets in a shared observation-aligned scene frame. SceneReGen addresses the generation-reconstruction gap through selective pose factorization: each object's observed orientation is encoded directly in the generated mesh, while translation and scale are estimated from instance-level and global scene evidence. Given a scene image and instance masks, a geometry encoder extracts dense cues; learnable shape queries condition a pretrained DiT-based 3D generator to produce complete meshes in their observed orientations, while position queries fuse object and scene features to assemble them in the shared frame. On the 3D-FUTURE evaluation subset, SceneReGen achieves the best scene-level CD, scene-level F-Score, and 3D bounding-box IoU among the evaluated methods, ties the best object-level CD, and ranks second in object-level F-Score. Qualitative outputs in autonomous-driving and embodied-AI scenes further illustrate the potential of asset-centric reconstruction beyond indoor furniture.