Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
Stripped of safety, language models can be tricked into generating misleading yet confident responses, with up to 90% of outputs being decoys under attack.
One in twenty papers from top-tier conferences may contain multiple hallucinated citations, raising serious concerns about the integrity of the academic record.
AdvGRPO enables robust attacker-defender co-training that significantly improves defender performance on safety benchmarks while generating effective attacks.