Search papers, labs, and topics across Lattice.
This paper introduces E-Commerce Bench, an innovative open-source benchmark designed to evaluate Large Language Models (LLMs) on long-horizon autonomous business operations, encompassing tasks such as market research, supplier negotiation, and cash flow management over a simulated year. The benchmark utilizes real e-commerce data and incorporates dynamic events like promotions and supply-chain shocks to create a realistic operational environment for LLM agents. Evaluation of 18 advanced models reveals that while GPT-5.6 Sol achieves the highest total assets, it underperforms in fraud avoidance and operational efficiency compared to other models, highlighting the trade-offs in LLM performance across different business metrics.
No single LLM excels across all dimensions of long-horizon business operations, revealing critical trade-offs in performance metrics like asset growth and fraud avoidance.
Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environments and long-range dependencies require Large Language Models (LLMs) to continually explore, learn from experience, and adapt their policies over thousands of steps. We introduce E-Commerce Bench, the first open-source benchmark that integrates multi-round counterpart negotiation and dynamic events into a year-long business operation. Over a 365-day year, an LLM agent concurrently runs multiple online stores, researching the market, negotiating with suppliers to source inventory, optimizing sales strategies, fulfilling orders, handling returns, and managing cash flow to maximize its end-of-year total assets. To construct a realistic merchant-side operating environment, the product and supplier data are derived from a real e-commerce platform, while a year-long calendar of promotions, natural disasters, and supply-chain shocks continually reshapes demand. For reproducibility, both sides of the market are deterministic: customer purchases and returns follow a fixed demand model, while a negotiation kernel determines supplier pricing, concessions, and decisions, with an LLM used only to verbalize them. We evaluate 18 frontier models across seven dimensions, including year-end assets, and find that no single model dominates. GPT-5.6 Sol earns the most, growing the 100,000 opening stake into 1,431,425, yet it ranks 16th of 18 on fraud avoidance and trails Fable5 in operational efficiency. Among open-weight models, Qwen3.8-Max-Preview leads with 416,252, 38% above GLM 5.2 (high), and achieves the strongest learning over the horizon, progressively bargaining down prices across repeated orders. Our code is available at https://github.com/QwenLM/E-CommerceBench.