Search papers, labs, and topics across Lattice.
Affiliation:
4
0
5
Dataset representation in NLP is alarmingly uneven, with geographic blind spots that could skew AI development and policy.
Deductive stereotyping in LLMs leads to biased reasoning, but targeted injection phrases can dramatically enhance fairness across diverse tasks.
Misfired alignment in LLMs can lead to a 18.9% failure rate in reasoning about stereotypes, revealing a critical flaw in current safety-oriented training methods.
A single biased example can completely undermine the alignment of large language models, revealing a critical flaw in their post-training safety mechanisms.