collaborators

8 papers

cs.SD2026

Which Data Matter? Embedding-Based Data Selection for Speech Recognition

Zakaria Aldeneh, Skyler Seto, Maureen de Seyssel +8

Modern ASR systems are typically trained on large-scale pseudo-labeled, in-the-wild data spanning multiple domains. While such heterogeneous data benefit generalist models designed…

cs.CL2025

TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different Languages

Minsu Kim, Jee-weon Jung, Hyeongseop Rha +5

The capability to jointly process multi-modal information is becoming an essential task. However, the limited number of paired multi-modal data and the large computational requirem…

eess.AS2025

Text-To-Speech Synthesis In The Wild

Jee-weon Jung, Wangyou Zhang, Soumi Maiti +11

Traditional Text-to-Speech (TTS) systems rely on studio-quality speech recorded in controlled settings.a Recently, an effort known as noisy-TTS training has emerged, aiming to util…

cs.CL2025

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

Siddhant Arora, Jinchuan Tian, Hayato Futami +5

Unlike traditional cascaded pipelines, end-to-end (E2E) spoken dialogue systems preserve full differentiability and capture non-phonemic information, making them well-suited for mo…

eess.AS2025

Context-Driven Dynamic Pruning for Large Speech Foundation Models

Masao Someki, Shikhar Bharadwaj, Atharva Anand Joshi +7

Speech foundation models achieve strong generalization across languages and acoustic conditions, but require significant computational resources for inference. In the context of sp…

cs.SD2025

SpoofCeleb: Speech Deepfake Detection and SASV In The Wild

Jee-weon Jung, Yihan Wu, Xin Wang +11

This paper introduces SpoofCeleb, a dataset designed for Speech Deepfake Detection (SDD) and Spoofing-robust Automatic Speaker Verification (SASV), utilizing source data from real-…