Search papers, labs, and topics across Lattice.
Affiliation:
3
0
6
2
Reward hacking can be mitigated with a simple one-line fix that improves out-of-distribution performance while keeping training robust.
LLMs can't reliably predict scientific experiment outcomes, and more worryingly, they have no idea when they're wrong, unlike human experts whose accuracy skyrockets when they feel confident.
LALMs can be easily tricked into "hearing" things that aren't there, with success rates as high as 95% on targeted attacks.