Search papers, labs, and topics across Lattice.
This paper extends the Abductive Latent Explanations (ALE) framework to accommodate non-Euclidean prototype-based neural networks, addressing a significant limitation of existing ALE formulations that are restricted to Euclidean spaces. By deriving novel bounding algorithms tailored to various geometric representations, the authors provide a systematic approach to generate formal explanations that ensure predictive safety and human readability across diverse architectures. The validation of these theoretical constructs on fully trained image classifiers demonstrates the potential for rigorous cross-architecture comparisons in interpretability, enhancing the understanding of model behavior in complex settings.
Non-Euclidean architectures can now be interpreted rigorously, bridging a critical gap in explainability for modern neural networks.
Prototype-based neural networks are hailed as interpretable-by-design architectures. Recently, Abductive Latent Explanations (ALE) were introduced to provide formal, mathematically guaranteed explanations that leverage the intrinsic structure of these networks to ensure both predictive safety and human readability. ALEs rely on computing tight bounds on latent space distances to produce formal explanations. However, existing ALE formulations are rigidly confined to Euclidean latent spaces. This leaves a critical gap: modern state-of-the-art architectures increasingly rely on non-Euclidean representations - such as spherical metrics, Gaussian densities, and dimensional projections - rendering current formal explanation methods incompatible. In this work, we generalize the ALE framework to support non-Euclidean prototype architectures. For each geometric variant, we systematically derive how to either map the architecture to existing bounds or construct novel, architecture-specific bounding algorithms. We validate our theoretical constructions by computing subset-minimal formal explanations on fully trained image classifiers. By unifying these diverse models under a single formal framework, we enable the first rigorous, cross-architecture comparison of their interpretability.