TurinUniversity of Southern DenmarkMay 25, 2026arXiv:2605.26045

Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals

Federico Torrielli, Peter Schneider-Kamp, Lukas Galke Poech

AI Summary

This paper investigates the calibration of confidence scores for activation oracles, which aim to interpret language model internals by translating activations into natural language. They evaluate six different methods for estimating the confidence of activation oracles on 6,000 samples per oracle, varying verbalizers and context prompts. Results show that bootstrap mode frequency is the best-calibrated method (ECE 5.7% on Qwen3-8B), while log-probability can serve as a fast, cheaper triage signal.

Key Contribution

Activation oracles, meant to make LLM internals legible, often produce poorly calibrated confidence scores, but a simple bootstrap method can significantly improve reliability.

Abstract

Activation oracles aim to make the activations of other models legible to humans and yield promising results compared to white-box interpretability techniques. However, uncertainty quantification (UQ) for the natural-language outputs of such activation oracles is so far understudied. Here, we investigate 6 different methods for estimating the confidence of activation oracles and evaluate how well-calibrated their confidence scores are. Our experiments on 6,000 samples per oracle (varying verbalizer and context prompts) reveal that bootstrap mode frequency is the best-calibrated method among those tested (ECE 5.7% vs. 25.5% for the answer-word log-probability on Qwen3-8B; 10.3% vs. 13.1% on Qwen3.6-27B), and that the log-prob baseline can serve as a fast triage signal at a fraction of the cost. Code and the patched trainer are available at https://github.com/federicotorrielli/probabilistic_activation_oracles.

Eval Frameworks & Benchmarks Interpretability & Mechanistic Interp Natural Language Processing

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals

Related Papers