Search papers, labs, and topics across Lattice.
3
0
6
3
Truthfulness in NLP research has surged to 37% of papers by 2026, reflecting a critical shift in focus towards safety and alignment in generative systems.
Semantic watermarks, embedded via AMR, survive paraphrasing attacks that obliterate token-level watermarks.
Current red-teaming efforts miss the forest for the trees: ARES reveals that safety failures often stem from a systemic breakdown between the LLM *and* the reward model, not just the LLM itself.