Search papers, labs, and topics across Lattice.
Affiliation:
4
0
5
Where an LLM represents a stereotype is detached from where it acts on it: linear decodability peaks up to 53% of model depth earlier than causal attribution, while fewer than 18% of stereotype-associated SAE features transfer across languages.
Steering vectors for distinct model behaviors remain naturally orthogonal in the residual stream, enabling training-free, multi-attribute control over language, safety, and style simply by stratifying injection across model depth.
Translation in multilingual LLMs is more modular than previously thought, with syntax and surface language production operating as distinct processes.
Tokenization isn’t just a preprocessing step; it’s a hidden driver of model performance and learning dynamics that could redefine how we interpret model evaluations.