Search papers, labs, and topics across Lattice.
This paper introduces TongueReenact, a novel framework for transferring tongue dynamics in face reenactment, addressing the anatomical inconsistencies caused by previous methods that overlook tongue motion. By leveraging a foundation-model-assisted bootstrapping pipeline, the authors create a tongue segmentation model that operates effectively without the need for curated annotations. The proposed spatially constrained latent masked diffusion model significantly enhances tongue synthesis, achieving over two times improvement on tongue-specific metrics compared to existing baselines while demonstrating perceptual superiority through a new VLM-based evaluation protocol.
Achieving anatomically consistent mouth interiors in face reenactment by accurately synthesizing tongue dynamics could revolutionize the realism of virtual avatars.
Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tongue dynamics, leading to anatomically inconsistent mouth interiors during speech and expressive motions. We introduce the first framework for cross-identity tongue dynamics transfer in face reenactment. We propose a foundation-model-assisted bootstrapping pipeline that produces a dedicated tongue segmentation model for in-the-wild reenactment without curated annotations. We further introduce a spatially constrained latent masked diffusion model for realistic tongue synthesis, with adaptive mask dilation for seamless mouth boundary transitions. Extensive experiments demonstrate improvements of more than two times over all baselines on every tongue-specific metric. We additionally propose a VLM-based evaluation protocol that replicates expert annotation at scale, confirming perceptual superiority across all ablation variants.