ARTFEED — Contemporary Art Intelligence

DuplexWorld: A New Benchmark for Evaluating Voice Agents Across Diverse Real-World Scenarios

ai-technology · 2026-08-13

A new benchmark called DuplexWorld has been introduced to evaluate speech-to-speech (S2S) voice agents in a more holistic manner. The benchmark, detailed in a paper on arXiv (2608.10716), addresses the limitations of existing benchmarks that focus primarily on agentic tool calling against databases. DuplexWorld creates six distinct worlds—banking, insurance, travel, healthcare, logistics, and pathfinding—where voice agents are particularly useful. It evaluates agents across eleven types of conversations and 156 scenarios, totaling over 350 hours of conversation, testing both conversational and analytical capabilities. The authors argue that current benchmarks fail to account for the diversity of conversational dialogue in mundane activities and do not test how faithfully an agent can assist on tasks beyond database manipulation. DuplexWorld aims to fill this gap by providing a more comprehensive evaluation framework.

Key facts

  • DuplexWorld is a new benchmark for evaluating speech-to-speech (S2S) voice agents.
  • It was introduced in a paper on arXiv with ID 2608.10716.
  • The benchmark covers six worlds: banking, insurance, travel, healthcare, logistics, and pathfinding.
  • It includes eleven types of conversations across 156 scenarios.
  • The total conversation time exceeds 350 hours.
  • Existing benchmarks are criticized for focusing on agentic tool calling against databases.
  • DuplexWorld aims to test conversational and analytical capabilities in diverse real-world tasks.
  • The paper is a cross announcement, indicating it may have been presented elsewhere.

Entities

Institutions

  • arXiv

Sources