Search papers, labs, and topics across Lattice.
This paper introduces SACHA, a novel framework for compressing animatable 3D Gaussian head avatars by integrating semantic-aware density control and appearance-motion decomposition. By intelligently allocating Gaussian primitives based on the visual saliency of different head regions and reducing temporal redundancy, SACHA significantly enhances the efficiency of storage and transmission without sacrificing rendering quality. Experimental results show that SACHA outperforms existing methods in rate-distortion performance while ensuring high-fidelity novel-view rendering of head avatars.
SACHA achieves superior compression of 3D Gaussian head avatars by leveraging semantic awareness, resulting in both storage efficiency and high-quality rendering.
Animatable 3D Gaussian head avatars offer high-fidelity and flexible facial rendering, but typically require substantial storage and transmission costs for numerous Gaussian primitives. Existing Gaussian head avatar methods overlook the visual saliency of different head semantic regions for more appropriate Gaussian primitive allocation, as well as the efficient compression of trained head avatar sequences. To tackle this obstacle, we propose SACHA, a dynamic head avatar compression framework that leverages both semantic-aware density control and appearance-motion decomposition to achieve compact representation and high-quality novel-view rendering of head avatar sequences. Specifically, the semantic-aware density control guides the adaptive allocation of Gaussian primitives across different head regions with region-adaptive densification and pruning. In addition, the appearance-motion decomposed compression further reduces the temporal redundancy of the avatar sequence by transmitting only head-prior parameters for avatar movements. Together, these designs enable a compact representation for efficient transmission of dynamic Gaussian head avatars while preserving visual fidelity. Experiments demonstrate that SACHA achieves a superior rate-distortion performance over existing Gaussian head avatar representation and compression methods while maintaining high-quality novel-view and novel-expression rendering.