activity
20202026
most citedTime-domain speech super-resolution with GAN based modeling for telephony speaker verification

1 citations · 1 across the 8 of their papers we have counts for

collaborators

11 papers

cs.CL2026

When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

Yen-Ju Lu, Yuzhe Wang, Yaohan Guan +8

Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluat…

cs.AI2026

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

Yuzhe Wang, Thomas Thebaud, Jennifer Hu +5

Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceB…

cs.CL2026

Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech

Kaavya Chaparala, Thomas Thebaud, Jesús Villalba López +3

There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed systemic errors and bias int…

eess.AS2026

Layer-Aware Early Fusion of Acoustic and Linguistic Embeddings for Cognitive Status Classification

Krystof Novotny, Laureano Moro-Velázquez, Jiri Mekyska

Speech contains both acoustic and linguistic patterns that reflect cognitive decline, and therefore models describing only one domain cannot fully capture such complexity. This stu…

eess.AS2023

Self-FiLM: Conditioning GANs with self-supervised representations for bandwidth extension based speaker recognition

Saurabh Kataria, Jesús Villalba, Laureano Moro-Velázquez +2

Speech super-resolution/Bandwidth Extension (BWE) can improve downstream tasks like Automatic Speaker Verification (ASV). We introduce a simple novel technique called Self-FiLM to…

eess.AS2022★ 1 cited

Time-domain speech super-resolution with GAN based modeling for telephony speaker verification

Saurabh Kataria, Jesús Villalba, Laureano Moro-Velázquez +2

Automatic Speaker Verification (ASV) technology has become commonplace in virtual assistants. However, its performance suffers when there is a mismatch between the train and test d…