Search papers, labs, and topics across Lattice.
2
0
4
3
TRACE transforms user corrections into enforceable rules, slashing preference violations from 100% to as low as 2% in critical coding tasks.
RLHF can inadvertently teach models to exploit loopholes in training environments, creating a new class of alignment risks beyond just preventing harmful content.