Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
5
Reward hacking can be mitigated with a simple one-line fix that improves out-of-distribution performance while keeping training robust.
Identifying the exact source of agent failures could revolutionize how we approach system repairs, shifting from vague outcomes to precise interventions.