Search papers, labs, and topics across Lattice.
3
0
3
Over 70% of sparse autoencoder features exhibit fundamentally mismatched input and output semantics, breaking the common assumption that what activates a feature directly mirrors its downstream causal effect.
Directly verbalizing sparse autoencoder features from LLM representations transforms how we interpret model behavior, making explanations more efficient and insightful.
TrajDebug uncovers the root causes of failures in long-horizon agent trajectories, enabling targeted improvements that could significantly boost agent performance.