Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
3
The findings show that even a direction orthogonal to a concept at one layer can contribute to the concept's downstream amplification, and demonstrate that more fine-grained category residuals should also be considered beyond shared general harmfulness representation to fully understand LLM safety.
LLMs exhibit stark relational biases in humor, refusing jokes from marginalized speakers up to 67.5% more often and judging them as more malicious, revealing a hidden dimension of unfairness beyond simple stereotyping.