Search papers, labs, and topics across Lattice.
This paper introduces StartupBench, a novel end-to-end (E2E) benchmark designed to evaluate general-purpose AI agents based on real-world workflows from market-validated AI startup products. By analyzing actual user demands and translating them into deliverable-oriented tasks, the authors reveal that even the most advanced models only achieve about 30% success on these tasks, highlighting significant gaps in their capabilities. The findings underscore the need for benchmarks that reflect practical applications, as many workflows remain challenging for current AI systems to handle reliably.
Despite advances in AI, even top models struggle with real-world tasks, achieving only 30% success on a benchmark grounded in market-validated workflows.
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.