Search papers, labs, and topics across Lattice.
3
1
4
4
By making environment design a learnable process, SPADE unlocks a new frontier in self-improvement for language agents, leading to substantial performance gains across diverse tasks.
Relying on a single simulator in multi-agent RL leads to dangerous mode collapse, but innovative solutions can boost generalization and performance by up to 14%.
The hardest AI tasks remain largely unsolved, with current models achieving only a 2.6% success rate on economically valuable workflows.