Search papers, labs, and topics across Lattice.
Affiliation:
6
1
6
2
No single LLM excels across all dimensions of long-horizon business operations, revealing critical trade-offs in performance metrics like asset growth and fraud avoidance.
A unified assessment framework reveals hidden insights about agent performance, transforming how we evaluate AI systems.
The hardest AI tasks remain largely unsolved, with current models achieving only a 2.6% success rate on economically valuable workflows.
No single AI model dominates across all professional industries, revealing distinct occupational capability profiles and highlighting the need for specialized AI development.
Training web agents in a simulator can now match real-world performance: Qwen3-14B, fine-tuned with WebWorld-synthesized trajectories, rivals GPT-4o on WebArena.
ToolRMs drastically improve tool-use accuracy in LLMs, outperforming existing models by up to 17.94%, while also reducing output token usage by over 66% through efficient inference-time scaling.