6 papers · 1 filter
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
Thomas Thebaud, Yuzhe Wang, Hao Zhang +5
Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these sy…
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue
Hao Zhang, Thomas Thebaud, Georgi Tinchev +2
Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluations often rely on human judgments or behav…
TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech
Sathvik Manikantan Napa Ugandhar, Hao Zhang, Alison Gunzler +5
With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly important. Emotional entrainment is shaped by…
Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems
Ashish Hallur, Thomas Thebaud, Georgi Tinchev +2
Speech-to-speech (S2S) AI agents are advancing rapidly, yet evaluation lacks interpretable speech-native measures for conversational prosody and rhythm. Because , speaking rat…
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
Alexandrine Fortier, Thomas Thebaud, Jesús Villalba +3
Speech language models (SLMs) are systems of systems: independent components that unite to achieve a common goal. Despite their heterogeneous nature, SLMs are often studied end-to-…
Demographic Attributes Prediction from Speech Using WavLM Embeddings
Yuchen Yang, Thomas Thebaud, Najim Dehak
This paper introduces a general classifier based on WavLM features, to infer demographic characteristics, such as age, gender, native language, education, and country, from speech.…