Search papers, labs, and topics across Lattice.
This paper introduces SpiD, a novel framework for creating photorealistic animatable head avatars from a single image by employing dual-axis disentanglement. By internalizing per-frame driving and decomposing facial features into three specialized Gaussian branches, SpiD eliminates the need for external tracking and enhances both expressiveness and rendering fidelity. Experimental results show that SpiD outperforms state-of-the-art methods in terms of performance and achieves the fastest inference speed on a single GPU, even with the complete driving pipeline included.
Achieving photorealistic head avatars from a single image without external tracking, SpiD sets a new standard for speed and fidelity in digital human synthesis.
Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian Splatting methods have achieved promising results, they rely on external tracking pipelines whose latency is excluded from inference measurements. Furthermore, they adopt unified representations that entangle geometrically distinct facial regions, limiting both expressiveness and rendering fidelity. We propose SpiD (Split and Drive), a single-image Gaussian head avatar framework built on two disentanglement axes. The compute axis internalizes per-frame driving, eliminating external tracking dependency at inference. The feature axis decomposes the avatar into three specialized Gaussian branches, each modeling a geometrically distinct facial domain. Extensive experiments demonstrate consistently strong performance against state-of-the-art methods while achieving the fastest inference speed among all compared methods on a single GPU with the complete driving pipeline included.