← Back to feed
2026-08-10data

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols

PDF preview for Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
Read on arXiv →

Key claim

New TTS evaluation framework captures nuanced speech dimensions.

In plain English

Imagine you're developing a text-to-speech (TTS) system that needs to sound as natural as possible. Currently, most evaluations rely on Mean Opinion Scores (MOS) that focus heavily on overall sound quality, but this approach often misses the subtleties of what makes speech feel natural to listeners. For instance, it might overlook issues like intonation, rhythm, or emotional expression, which are crucial for a realistic experience. This limitation is what's called a collapse onto acoustic signal quality, meaning it doesn't capture the full spectrum of human perception in speech. Additionally, newer methods using Audio Large Language Models (Audio-LLMs) show promise but struggle with consistency across different speech dimensions, leading to selective and prompt-dependent evaluations. This is where the new approach comes in. By breaking down 'naturalness' into ten distinct perceptual dimensions, the authors created a comprehensive evaluation benchmark that includes 860 utterances rated by trained linguists. This allows for a more nuanced understanding of TTS quality, revealing that existing methods often fail to address the breadth of linguistically structured speech errors. With this new dataset and evaluation framework, builders can better assess and improve TTS systems, leading to more natural-sounding speech outputs.

Novelty
8.0/10

The paper introduces a new evaluation framework for TTS that dissects naturalness into distinct dimensions.

Reliability
7.5/10

The benchmarking results are based on a well-annotated dataset and rigorous evaluation of existing methods.

Deep reliability assessment

The methodology supports a detailed analysis of TTS evaluation by breaking down naturalness into 10 dimensions, but it may overclaim by suggesting it fully captures human perceptual judgment across all dimensions.

Reproducibility

Yes, the dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.

Key figure

Figure 1 illustrates the dataset construction pipeline, showing how sentences are processed through error-generation pathways and annotated by linguists.