Search papers, labs, and topics across Lattice.
This paper introduces a novel transfer framework for the high-fidelity removal of eyeglasses from video, addressing challenges posed by refractive distortions and specular reflections. By utilizing a three-stage structural filtering process and physically-based simulation of lens optics, the approach maintains identity, expression, and pose while generating diverse, paired data for training. Evaluations demonstrate that the proposed JFSnet architecture significantly outperforms existing diffusion and GAN-based methods in terms of ocular consistency, temporal stability, and restoration quality, achieving an inference speed of 27.68 FPS.
Eyeglasses removal from video can now be achieved with unprecedented fidelity and temporal stability, thanks to a physics-grounded approach that preserves identity and expression.
High-fidelity removal of eyeglasses from video is a major challenge in facial attribute editing, as the underlying facial geometry is often obscured by complex refractive distortions and view-dependent specular reflections. While large-scale generative priors have shown promise in eye-glasses removal via static image inpainting, they often lack the structural constraints necessary to maintain identity, expression, and pose, leading to visible"identity drift"in both static images and dynamic sequences. In this paper, we propose a novel transfer framework that addresses the stochastic nature of generative priors. Our pipeline first extracts high-fidelity synthetic face images from a commercial-grade generative model (Nano Banana, Gemini 3 Pro Image), regularizes them via a three-stage structural filtering process to preserve identity, expression, and pose, and finally applies physically-based simulation of lens optics during training to provide diverse, paired data. This process transfers Nano Banana's photo-realistic, multi-view knowledge into a specialized restoration architecture, JFSnet (Joint Feature-Spatial network). JFSnet integrates DINOv2-based semantic features with a convolutional decoder for spatial reconstruction, leveraging translation equivariance constraints to improve temporal consistency and high-frequency detail preservation. Evaluations on the curated Flickr-Faces-HQ (FFHQ) subset (12,163 images) show that our approach achieves high fidelity and structural accuracy, while maintaining inference speed of 27.68 FPS. In perceptual studies on CelebV-Text video sequences, our results are consistently preferred over diffusion and GAN-based baselines for ocular consistency, temporal stability, and overall restoration quality.