18 citations · 59 across the 53 of their papers we have counts for
12 papers · 1 filter
Who Are They to Each Other? Multi-Agent Reasoning for Speaker Relationship Inference
Yaohan Guan, Yen-Ju Lu, Yuzhe Wang +5
Inferring speaker relationships from spoken conversations is an important step towards socially aware speech understanding. However, this task remains underexplored, and supervised…
When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue
Yen-Ju Lu, Yuzhe Wang, Yaohan Guan +8
Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluat…
Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection
Mingrui Liang, Thomas Thebaud, Lukasz Wojciak +4
Recent advances in speech synthesis and voice conversion, which pose threats to security and privacy, have underscored the need for deepfake detection technology. Although existing…
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
Thomas Thebaud, Junhyeok Lee, Laureano Moro-Velazquez +2
Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather t…
StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech
Yuzhe Wang, Thomas Thebaud, Jennifer Hu +5
Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceB…
Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System
Thomas Thebaud, Sonal Joshi, Henry Li +4
Poisoning attacks entail attackers intentionally tampering with training data. In this paper, we consider a dirty-label poisoning attack scenario on a speech commands classificatio…