Search papers, labs, and topics across Lattice.
This paper introduces AvatarDynamizer, a generative approach that enhances static 3D avatars by embedding realistic surface dynamics and enabling multi-view consistency. By leveraging a novel texture-space surface-dynamics embedding and formulating avatar dynamics as conditional texture generation, the method allows for the transformation of off-the-shelf avatars into dynamic, controllable 4D representations. Experimental results demonstrate that AvatarDynamizer significantly outperforms existing methods in visual fidelity, particularly when trained on limited dynamic data, addressing a critical gap in realistic avatar animation.
Transforming static avatars into dynamic, realistic representations could redefine the standards for avatar realism in virtual environments.
For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism. Person-agnostic methods recover static 3D avatars from monocular images, videos, or text prompts, but their skeleton-driven animations lack realistic surface dynamics such as clothing wrinkles. In contrast, person-specific methods achieve high-quality rendering and realistic dynamics, but require expensive multi-view captures for each individual. Recent generalizable dynamic avatar methods struggle to embed surface dynamics, leading to either limited multi-view consistency or dynamic expressiveness. To this end, we propose AvatarDynamizer, a generative method that transforms an off-the-shelf static 3D avatar into a controllable, realistic, and multi-view-consistent 4D avatar. We introduce a novel texture-space surface-dynamics embedding and formulate avatar dynamics modeling as conditional texture generation. Our encoder--decoder representation embeds pose-dependent dynamics into dynamic texture maps, enabling compatibility with pre-trained video diffusion models while decoding them into 3D Gaussians for multi-view consistent rendering. Since existing datasets are limited in scale, sequence length, or motion diversity, we collect a large-scale multi-view dataset with long sequences covering diverse skeletal motions and surface dynamics. Experiments show that our method effectively animates static avatars with faithful surface dynamics and outperforms competing generalizable methods in visual fidelity, especially under limited dynamic training data.