Search papers, labs, and topics across Lattice.
The University of Texas at Austin
4
5
6
4
Induced anger can lock LLMs into poor decision patterns by reducing their sensitivity to penalties, unlike human decision-making.
Existing Multi-modality Machine Unlearning methods are fundamentally flawed, allowing adversaries to recover nearly all supposedly erased sensitive information from MLLMs.
Traditional research papers are costing AI agents reproducibility and understanding, but a new "Agent-Native" format that captures the full messy research process boosts performance by up to 20%.
Using preference data from stronger models to align LLMs via DPO can backfire, dramatically worsening safety by making models more susceptible to jailbreaking.