Search papers, labs, and topics across Lattice.
This paper introduces Phoneme-Driven Gaussian Splatting (PD-GS), a novel approach that enhances 3D Gaussian Splatting for audio-driven talking heads by integrating time-aligned phoneme tokens to improve lip articulation accuracy. The method addresses the common issue of over-smoothed mouth movements and closure violations by using a Linguistic Fusion Module (LFM) that combines continuous audio context with discrete phoneme embeddings. Experimental results on the HDTF dataset show that PD-GS achieves superior lip geometry and significantly reduces closure violations compared to existing baselines, leading to more linguistically accurate neural avatars.
Lip articulation accuracy improves dramatically when phoneme information is integrated into audio-driven rendering, reducing closure violations and enhancing realism.
3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth motion is often over-smoothed and may violate hard articulatory constraints such as bilabial closures, producing the notorious ``leaky mouth'' artifact. A key difficulty is that brief, discrete articulatory events are inferred from a continuous acoustic embedding under a regression objective, which biases predictions toward averaged mouth configurations. While modern self-supervised speech encoders provide rich prosodic and phonetic cues, they do not provide an explicit, frame-aligned linguistic target that reliably disambiguates closure-level events. We propose \textbf{Phoneme-Driven Gaussian Splatting (PD-GS)}, which augments a 3DGS talker with time-aligned phoneme tokens obtained from an automatic ASR and forced-alignment pipeline. Our core component, the \textbf{Linguistic Fusion Module (LFM)}, adaptively fuses continuous audio context with discrete phoneme embeddings through a learned gate, allowing the model to preserve smooth audio-driven dynamics while strengthening phoneme guidance on articulation-critical segments. PD-GS is trained purely from monocular video using image reconstruction and lip landmark supervision. On HDTF, PD-GS achieves the best lip geometry among the compared baselines (LMD 2.66) and qualitatively reduces closure violations in challenging phoneme sequences, yielding more linguistically faithful neural avatars.