activity
20182026
most citedNon-Contrastive Self-supervised Learning for Utterance-Level Information Extraction from Speech

19 citations · 72 across the 37 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

Yen-Ju Lu, Yuzhe Wang, Yaohan Guan +8

Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluat…

cs.CL2025

Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization

Yen-Ju Lu, Kunxiao Gao, Mingrui Liang +5

Recent audio language models can follow long conversations. However, research on emotion-aware or spoken dialogue summarization is constrained by the lack of data that links speech…

cs.CL2025

Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models

Alexandrine Fortier, Thomas Thebaud, Jesús Villalba +3

Speech language models (SLMs) are systems of systems: independent components that unite to achieve a common goal. Despite their heterogeneous nature, SLMs are often studied end-to-…

cs.CL2025

Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation

Yen-Ju Lu, Thomas Thebaud, Laureano Moro-Velazquez +2

We present Paired by the Teacher (PbT), a two-stage teacher-student pipeline that synthesizes accurate input-output pairs without human labels or parallel data. In many low-resourc…

cs.CL2021

Beyond Isolated Utterances: Conversational Emotion Recognition

Raghavendra Pappagari, Piotr Żelasko, Jesús Villalba +2

Speech emotion recognition is the task of recognizing the speaker's emotional state given a recording of their utterance. While most of the current approaches focus on inferring em…

cs.CL2019

Hierarchical Transformers for Long Document Classification

Raghavendra Pappagari, Piotr Żelasko, Jesús Villalba +2

BERT, which stands for Bidirectional Encoder Representations from Transformers, is a recently introduced language representation model based upon the transfer learning paradigm. We…