Search papers, labs, and topics across Lattice.
2
0
3
9
Defensive poisoning can effectively clear over half of original backdoors in LLMs, but the dynamics of trigger recognition reveal deeper vulnerabilities.
Conditioning LLMs on human privacy judgments leads to a remarkable increase in alignment with user expectations, showcasing a new standard for agent training.