Search papers, labs, and topics across Lattice.
4
0
5
Qwen-AgentWorld achieves unprecedented simulation fidelity, outperforming existing models and enabling scalable agentic reinforcement learning across diverse real-world environments.
Task success rates for agentic phone use soar from 36.67% to 45.33% through a novel combination of real and mock environments in training.
Coding agents struggle to create complete and engaging games, with top performers barely reaching 41.46% success in end-to-end game generation.
Forget hand-crafted benchmarks: CUA-Gym's auto-generated training data lets computer-use agents crush existing open-source models on real-world tasks.