Search papers, labs, and topics across Lattice.
Rutgers University
2
0
4
0
Sparse autoencoders reveal that a Transformer agent's decision-making strategies can be interpreted without explicit rule classification, uncovering hidden concepts that drive behavior.
By surgically intervening in MLLM decoding, this work cuts hallucination rates without sacrificing descriptive quality, a feat prior methods struggled to achieve.