Search papers, labs, and topics across Lattice.
This paper introduces HIVE-3D, a hierarchical voxel enhancement framework that addresses the limitations of existing methods in generating high-quality 3D scenes from single images. By employing image segmentation and attention-based retrieval, the method aligns 2D image components with 3D scene components, organizing them into a hierarchical component tree for refined voxel generation. Experimental results show that HIVE-3D achieves state-of-the-art performance in 3D scene generation, significantly surpassing previous approaches in quality and resolution.
HIVE-3D transforms single-image inputs into high-resolution 3D scenes, setting a new benchmark in quality and detail.
Recently, a line of works can generate impressive 3D objects from a single image, but they are limited by restricted representation resolution, making them unsuitable for 3D scene generation. In this work, we introduce HIVE-3D, a novel method for high-quality 3D scene generation based on hierarchical voxel enhancement framework. Specifically, given a single scene image as input, we first produce a coarse initial scene, then introduce image segmentation and attention-based retrieval to align 2D image components with 3D scene components. Subsequently, we organize these scene relations into a hierarchical component tree, where nodes closer to the leaves denote finer-grained components. Finally, we propose a voxel super-resolution model that generates refined voxels for the target instance while maintaining strong consistency with the coarse voxels. Equipped with this model, we perform coarse-to-fine hierarchical super-resolution on images and voxels for each component, producing a high-resolution and high-quality 3D scene. Extensive experiments demonstrate that our method significantly outperforms previous approaches, achieving state-of-the-art performance.