Search papers, labs, and topics across Lattice.
Independent
3
0
4
Runtime safety detection in coding agents can be achieved through a novel intervention in hidden representations, drastically reducing harmful actions during multi-turn interactions.
Structural differences in circuits may be misleading, as they often reflect interchangeable mechanisms rather than distinct functionalities.
LLMs respond to increasingly difficult out-of-distribution inputs by activating sparser representations in their last hidden states, revealing a quantifiable relationship between task difficulty and neural activity.