Search papers, labs, and topics across Lattice.
This paper introduces Axolotl3D, a novel framework for 3D shape completion that integrates multi-modal inputs, including images, visibility masks, camera parameters, and partial point clouds. By utilizing a unified training strategy that synthesizes diverse conditioning signals from extensive 3D datasets, Axolotl3D achieves robust performance in both clean and occluded scenarios. The model not only excels in state-of-the-art shape completion but also facilitates real-world reconstruction and geometry-consistent editing, addressing limitations of existing methods that operate under restrictive assumptions.
Axolotl3D achieves state-of-the-art 3D shape completion by leveraging multi-modal inputs, enabling robust performance even in challenging occluded settings.
Recent 3D generative models produce high-quality geometry from a single image using large-scale priors and diffusion architectures. However, they assume complete visibility and single-view inputs, limiting applicability in multi-view, occluded, or editing scenarios. Although prior works address these challenges individually, they lack a unified framework for controllable 3D completion under diverse conditioning signals. We present Axolotl3D, a multi-modal and occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud. The point cloud serves as a geometric anchor promoting faithful shape completion, while camera parameters ensure consistent multi-view alignment in a shared 3D coordinate system. A unified training strategy synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning. Experiments on Toys4K and OmniObject3D demonstrate state-of-the-art performance under both clean and occluded settings, as well as strong results in real-world reconstruction and geometry-consistent editing.