Apr 8, 2026arXiv:2604.07054

Sell More, Play Less: Benchmarking LLM Realistic Selling Skill

Xuanbo Su, Wenhao Hu, Le Zhan, Yanqi Yang, Leo Huang

AI Summary

The authors introduce SalesLLM, a new bilingual benchmark for evaluating LLMs in realistic sales dialogues, focusing on deal progression and outcomes. They create a dataset of 1,805 multi-turn scenarios with controllable difficulty and personas, derived from financial services and consumer goods applications. To improve simulation fidelity, they train a user model, CustomerLM, using SFT and DPO on crowdworker-involved sales conversations, significantly reducing role inversion.

Key Contribution

LLMs can now be rigorously benchmarked on realistic sales skills, revealing a wide performance gap where some models rival humans while others fall short.

Abstract

Sales dialogues require multi-turn, goal-directed persuasion under asymmetric incentives, which makes them a challenging setting for large language models (LLMs). Yet existing dialogue benchmarks rarely measure deal progression and outcomes. We introduce SalesLLM, a bilingual (ZH/EN) benchmark derived from realistic applications covering Financial Services and Consumer Goods, built from 30,074 scripted configurations and 1,805 curated multi-turn scenarios with controllable difficulty and personas. We propose a fully automatic evaluation pipeline that combines (i) an LLM-based rater for sales-process progress, and (ii) fine-tuned BERT classifiers for end-of-dialogue buying intent. To improve simulation fidelity, we train a user model, CustomerLM, with SFT and DPO on 8,000 crowdworker-involved sales conversations, reducing role inversion from 17.44% (GPT-4o) to 8.8%. SalesLLM scores correlate strongly with expert human ratings (Pearson r=0.98). Experiments across 15 mainstream LLMs reveal substantial variability: top-performance LLMs are competitive with human-level performance while the less capable ones are worse than human. SalesLLM serves as a scalable benchmark for developing and evaluating outcome-oriented sales agents.

Data Curation & Synthetic Data Eval Frameworks & Benchmarks Natural Language Processing

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Sell More, Play Less: Benchmarking LLM Realistic Selling Skill

Related Papers