Search papers, labs, and topics across Lattice.
Nanyang Technological University, Agency for Science, Technology, and Research (A*STAR)
2
0
4
Agents can discover tool behaviors but fail to adapt effectively, often resorting to inefficient exhaustive searches in dynamic environments.
GRAIL reweights token advantages in reinforcement learning, leading to significant accuracy gains without the need for expensive process-level supervision.