Search papers, labs, and topics across Lattice.
2
0
2
Despite advances in AI, even top models struggle with real-world tasks, achieving only 30% success on a benchmark grounded in market-validated workflows.
Current AI agents struggle with long-horizon professional tasks, achieving only 30% success in complex GUI workflows, revealing critical gaps in their capabilities.