Search papers, labs, and topics across Lattice.
Saarland University, German Research Center for Artificial Intelligence (DFKI
1
0
2
Where an LLM represents a stereotype is detached from where it acts on it: linear decodability peaks up to 53% of model depth earlier than causal attribution, while fewer than 18% of stereotype-associated SAE features transfer across languages.