Search papers, labs, and topics across Lattice.
Affiliation:
3
0
6
0
Reward hacking can be mitigated with a simple one-line fix that improves out-of-distribution performance while keeping training robust.
Current video editing AIs still struggle to balance visual quality, instruction adherence, and localized edits, as revealed by a new benchmark designed to disentangle these factors.
LLM-generated explanations often fail to help users identify incorrect answers, and simply scaling models or applying post-training doesn't fix the problem.