Search papers, labs, and topics across Lattice.
2
0
5
0
Constraining rollout updates to complementary subspaces can enhance performance and stability in large language models, achieving up to 27.69 points improvement over existing methods.
Agent-repair leaderboards are more fragile than we thought: methods that peek at the evaluator's signals to guide internal repair choices can cause drastic reordering when the evaluator changes.