collaborators

5 papers

cs.CL2026

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Thomas Thebaud, Yuzhe Wang, Hao Zhang +5

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these sy…

cs.CL2026

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

Hao Zhang, Thomas Thebaud, Georgi Tinchev +2

Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluations often rely on human judgments or behav…

cs.CL2026

TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

Sathvik Manikantan Napa Ugandhar, Hao Zhang, Alison Gunzler +5

With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly important. Emotional entrainment is shaped by…

cs.CL2026

Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems

Ashish Hallur, Thomas Thebaud, Georgi Tinchev +2

Speech-to-speech (S2S) AI agents are advancing rapidly, yet evaluation lacks interpretable speech-native measures for conversational prosody and rhythm. Because , speaking rat…

cs.AI2026

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

Yuzhe Wang, Thomas Thebaud, Jennifer Hu +5

Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceB…