Search papers, labs, and topics across Lattice.
2
0
6
0
Decoupling task synthesis from on-policy rollouts solves the learning-signal saturation bottleneck in terminal agents, boosting long-horizon RL performance by up to 18 percentage points on Terminal-Bench 2.1.
Performance gaps in multilingual medical evaluations reveal that proprietary models outperform open-source ones, but translation quality can swing results dramatically.