JD.comApr 16, 2026arXiv:2604.15037

From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench

AI Summary

The paper introduces ProVoice-Bench, a new benchmark designed to evaluate the proactivity of voice agents across four novel tasks requiring proactive intervention and monitoring. They use a multi-stage data synthesis pipeline to generate 1,182 high-quality samples. Evaluation of current multimodal LLMs on ProVoice-Bench reveals significant performance gaps, especially in over-triggering and reasoning, highlighting areas for improvement in proactive agent design.

Key Contribution

Today's best multimodal LLMs are surprisingly bad at knowing *when* to speak up, revealing a critical flaw in the shift towards proactive voice agents.

Abstract

Recent advancements in LLM agents are gradually shifting from reactive, text-based paradigms toward proactive, multimodal interaction. However, existing benchmarks primarily focus on reactive responses, overlooking the complexities of proactive intervention and monitoring. To bridge this gap, we introduce ProVoice-Bench, the first evaluation framework specifically designed for proactive voice agents, featuring four novel tasks. By leveraging a multi-stage data synthesis pipeline, we curate 1,182 high-quality samples for rigorous testing. Our evaluation of state-of-the-art Multimodal LLMs reveals a significant performance gap, particularly regarding over-triggering and reasoning capabilities. These findings highlight the limitations of current models and offer a roadmap for developing more natural, context-aware proactive agents.

Eval Frameworks & Benchmarks Speech & Audio Tool Use & Agents

Citation Metrics

Citations0

Influential citations0

References28

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench

Related Papers