← Back to feed
2026-07-06agentsmultimodalcode

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar, Ashish Hallur, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Velazquez

PDF preview for SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
Read on arXiv →

Key claim

Current models struggle with natural conversational behavior despite high quality.

In plain English

Imagine you're building a system that can take spoken questions and respond with synthetic speech. You want it to sound natural, like a real conversation, but current benchmarks mainly focus on how accurately the system understands and generates speech. This is where things can go wrong: while a model might produce clear and correct answers, it could still feel robotic or awkward in a back-and-forth dialogue. For instance, it might interrupt too often, take too long to respond, or fail to adapt its tone to the emotional context of the conversation. These issues are what's called naturalness failures in conversational AI.

To address these shortcomings, the authors created SPEARBench, a new benchmark specifically designed to evaluate how naturally speech-to-speech models interact in conversations. Instead of just measuring accuracy, SPEARBench looks at various factors like response timing, emotional tone, and how well the model maintains consistency in language and dialect. By using controlled dialogue prompts and comparing model outputs to human responses, they provide a more comprehensive view of conversational quality.

What sets this work apart from previous benchmarks is its multidimensional approach to evaluation. It shows that even when models achieve high technical performance, they can still fall short in mimicking human conversational behavior. For anyone building conversational systems, this means you need to focus not just on getting the right answers, but also on how those answers are delivered to ensure a more human-like interaction.

Novelty
8.0/10

The introduction of SPEARBench provides a new framework for evaluating naturalness in speech-to-speech models, extending existing benchmarks.

Reliability
7.5/10

The evaluation includes multiple models and dimensions, but may lack extensive baselines for all claims.

Deep reliability assessment

The methodology supports evaluating naturalness in speech-to-speech models using a multidimensional protocol, but it may overclaim by not fully addressing the complexity of human conversational dynamics in diverse real-world settings.

Reproducibility

Yes, the dataset extracted from Seamless Interaction is available for download, and the code to apply the protocol to new evaluation datasets is provided.

Key figure

Figure 1 likely illustrates the architecture or evaluation pipeline of SPEARBench, showing how different metrics are integrated to assess conversational naturalness.

Codelink
arxiv.org/abs/2606.30543Official
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models — Frontier Papers