Search papers, labs, and topics across Lattice.
3
1
5
40
Externally generated explanations can rival human-written ones in boosting model accuracy, but the selection strategy for self-generated explanations can make or break their utility.
Reward design is a game-changer in reinforcement unlearning, enabling models to forget knowledge up to three times faster without sacrificing performance.
Hidden safety risks in AI systems may be more dangerous than the visible failures we obsess over, revealing a critical need for a new framework to diagnose them.