Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
3
Instruction-tuned LLMs can outperform specialized models in hate speech detection, achieving state-of-the-art results across multiple domains and languages.
CLA scores based on English can predict translation quality better than direct source-target alignment, highlighting English's role as a crucial pivot in multilingual LLMs.
LLMs consistently invert human judgments on evaluative attributes of hate speech (e.g., respect, sentiment), even while aligning on behavioral aspects, suggesting a fundamental misalignment in subjective understanding.