Search papers, labs, and topics across Lattice.
Rutgers University
1
0
1
23
Sparse autoencoders reveal that a Transformer agent's decision-making strategies can be interpreted without explicit rule classification, uncovering hidden concepts that drive behavior.