Search papers, labs, and topics across Lattice.
2
0
4
1
RAMP uncovers that agentic models can lose up to 80% of their effectiveness in complex, real-world workflows, a stark contrast to their performance in isolated benchmarks.
Current mobile GUI agents are surprisingly inept at everyday smartphone tasks, achieving only 62% success on a new benchmark of real-world Android apps.