Search papers, labs, and topics across Lattice.
9
0
13
8
Mobile agents can now navigate complex GUIs with unprecedented efficiency, thanks to a novel data-environment co-scaling framework.
Task success rates for agentic phone use soar from 36.67% to 45.33% through a novel combination of real and mock environments in training.
Coding agents struggle to create complete and engaging games, with top performers barely reaching 41.46% success in end-to-end game generation.
Reliable phone automation hinges on mixed-action capabilities, with agents achieving a 75% success rate in real-world workflows.
Forget hand-crafting mobile benchmarks – PhoneWorld lets you automatically generate them from real-world GUI trajectories, leading to massive performance gains for phone-use agents.
Long-context LLM rankings dramatically reshuffle when evaluated across a range of context lengths and capabilities, proving that a single headline score is misleading.
LLM agents still fail to reliably automate real-world workflows, with even the best models succeeding on only two-thirds of tasks in a new live benchmark.
Pruning reasoning paths with a learned "STOP" token slashes compute costs and boosts accuracy in large reasoning models, outperforming existing methods.
Current phone-use agents are often *too* helpful, routinely violating user privacy by filling in unnecessary personal information even when a task doesn't require it.