Search papers, labs, and topics across Lattice.
This paper introduces a novel approach to data attribution in machine learning for space missions, specifically focusing on the Ariel mission, by utilizing influence functions to assess training data impact on predictions. The authors reformulate influence in terms of prediction rather than loss, allowing for label-free deployment, and derive an influence-based error proxy that effectively correlates with spectral errors in simulated data. Key findings indicate that this method not only identifies influential training samples but also approximates harmful ones, establishing a robust operational framework for scientific machine learning applications.
Influence functions can now pinpoint the most impactful training samples without requiring labels, revolutionizing data attribution in scientific missions.
Interpretability is critical for machine learning models deployed in scientific space missions such as ESA's Ariel, where ground truth is unavailable during operations and physical plausibility must be assessed. While most explainable AI methods focus on feature attribution, this work investigates training data attribution through influence functions and introduces three key contributions for operational spectroscopy pipelines. First, influence is reformulated in terms of prediction rather than loss, enabling label-free deployment. Second, by leveraging the closed-form ridge solution of an Extreme Learning Machine, infinitesimal prediction influence is efficiently computed. Third, an influence-based conservative error proxy is derived by propagating training residuals through the influence sensitivities. Evaluated against simulated spectra, the proposed proxy correlates strongly with scale and shape-based spectral errors. Furthermore, influence functions enable the identification of the most influential samples and the approximation of the most harmful ones. Together, these results suggest that this approach can serve as an operational framework for scientific machine learning.