Search papers, labs, and topics across Lattice.
VoxStruct3D introduces a voxel-space flow-matching framework for high-fidelity 3D MRI synthesis that addresses the limitations of latent diffusion's reconstruction bottleneck. By employing a Structure-First, Image-Follows strategy, it integrates anatomical priors with advanced voxel generation techniques, enabling the model to generate anatomically coherent and visually realistic MRI volumes. Experimental results demonstrate that VoxStruct3D outperforms existing methods in feature-distribution alignment, sample diversity, and perceptual quality across various MRI datasets.
Achieving superior anatomical coherence and visual realism in 3D MRI synthesis, VoxStruct3D redefines the landscape of volumetric generation.
High-fidelity 3D MRI synthesis requires both globally coherent anatomy and fine-grained voxel-level detail. Although latent diffusion makes volumetric generation tractable, its image autoencoder introduces a reconstruction bottleneck that can limit the fine detail recoverable in the final volume. We present VoxStruct3D, a voxel-space flow-matching framework that directly models full-resolution MRI volumes using a clean-data prediction objective. Its Volumetric Voxel Generator (VVG) combines factorized 3D patch embedding with overlapping upsampling, time-modulated residual refinement, and skip fusion, enabling neighboring tokens to jointly reconstruct shared voxel regions and suppress patch-boundary artifacts. To complement direct voxel-space modeling with an explicit anatomical prior, we further introduce a Structure-First, Image-Follows (SFIF) strategy. A frozen pretrained 3D medical encoder and a StructVAE extract compact structure tokens that preserve dominant anatomy, while a structure-leading schedule keeps their trajectory ahead of the image trajectory. Patch-Aligned RoPE spatially aligns the unequal token grids, and asymmetric attention enforces one-way guidance from structure to image. Experiments on pathological and healthy T1-weighted brain MRI datasets show that VoxStruct3D achieves the strongest overall performance across feature-distribution alignment, sample diversity, and perceptual quality, producing anatomically coherent and visually realistic volumes.