collaborators

19 papers

eess.AS2026

Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection

Mingrui Liang, Thomas Thebaud, Lukasz Wojciak +4

Recent advances in speech synthesis and voice conversion, which pose threats to security and privacy, have underscored the need for deepfake detection technology. Although existing…

cs.CL2026

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Thomas Thebaud, Yuzhe Wang, Hao Zhang +5

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these sy…

eess.AS2026

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions

Thomas Thebaud, Junhyeok Lee, Laureano Moro-Velazquez +2

Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather t…

cs.CL2026

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

Hao Zhang, Thomas Thebaud, Georgi Tinchev +2

Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluations often rely on human judgments or behav…

cs.CL2026

TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

Sathvik Manikantan Napa Ugandhar, Hao Zhang, Alison Gunzler +5

With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly important. Emotional entrainment is shaped by…

cs.CL2026

Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems

Ashish Hallur, Thomas Thebaud, Georgi Tinchev +2

Speech-to-speech (S2S) AI agents are advancing rapidly, yet evaluation lacks interpretable speech-native measures for conversational prosody and rhythm. Because , speaking rat…