Search papers, labs, and topics across Lattice.
This study applies sparse autoencoders for mechanistic interpretability in a neutrino foundation model trained on IceCube data, revealing an atlas of physical concepts within the model's representation. Through rigorous validation and causal interventions, it was found that the model's direction head minimally utilizes this atlas, while an uncertainty head trained on the same representation significantly improves angular reconstruction accuracy. The interpretable estimator enhances median angular resolution from 20.2掳 to 3.2掳 at 20% selection efficiency, demonstrating the potential of mechanistic interpretability in extracting useful latent physics for practical applications.
Uncovering a rich atlas of physical concepts in a neutrino model reveals that interpretability can dramatically enhance angular resolution in particle physics tasks.
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on the same event-level representation to predict the model's angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At $20\%$ selection efficiency, this interpretable estimator improves the median angular resolution from $20.2^\circ$ to $3.2^\circ$. These results suggest that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.