Search papers, labs, and topics across Lattice.
University of Oxford
3
0
2
Slow features in Persistent SAEs can maintain detection signals over long contexts, revolutionizing how we interpret and monitor language models.
A unified theory of interpretability that leverages Lagrangian mechanics to transform opaque models into interpretable ones, revealing new research avenues and design principles.
Embracing information leakage in concept-based models can enhance accuracy and interpretability, turning a perceived flaw into a powerful asset.