Search papers, labs, and topics across Lattice.
Affiliation:
2
0
2
Neural networks exhibit an internal "feeling of error"鈥攎anifesting as sharply suppressed concept activations inside Sparse Autoencoders鈥攖hat reliably flags segmentation failures where standard output uncertainty metrics fail.
ViT weights harbor fully legible, input-invariant concept graphs of the world, unlocking targeted circuit interventions that improve shortcut removal by 11.0% over prior baselines.