Search papers, labs, and topics across Lattice.
This paper introduces a novel approach to mitigate catastrophic forgetting in fake speech detection by employing domain translators within a frozen detector framework. The proposed traceback translator network effectively remaps new feature spaces to original ones, allowing the model to adapt to new data while maintaining performance on previously encountered samples. Experimental results demonstrate that this method achieves superior detection rates compared to traditional retraining methods, all while reducing computational demands and preserving accuracy on older data.
Forgetting-resilient fake speech detection can be achieved without retraining, preserving accuracy and cutting computational costs.
Fake speech detectors are increasingly challenged by the development of new and more accurate generative models. To cope with this problem, continual learning techniques are nowadays widely considered feasible strategies for updating models to new datasets, but they also lead to decreased performance on previously seen samples (catastrophic forgetting). In this work, we propose a forgetting-resilient solution based on the adoption of domain translators within a frozen detector, which remaps the new feature spaces into the original ones by means of a traceback translator network. Experimental results show that this strategy enables the achievement of high detection rates with respect to traditional retraining, while minimizing the computational effort and preserving the detection accuracy on previous data.