activity
20242026
collaborators

5 papers

eess.AS2026

MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables

Sung-Lin Yeh, Wei Zhou, Gil Keren +6

Recent speech language models rely on encoders that are optimized separately from autoregressive models. Since these encoders are unaware of the downstream objectives, the extracte…

eess.AS2025

Learning Speech Representations with Variational Predictive Coding

Sung-Lin Yeh, Peter Bell, Hao Tang

Despite being the best known objective for learning speech representations, the HuBERT objective has not been further developed and improved. We argue that it is the lack of an und…

eess.AS2025

Whisper Has an Internal Word Aligner

Sung-Lin Yeh, Yen Meng, Hao Tang

There is an increasing interest in obtaining accurate word-level timestamps from strong automatic speech recognizers, in particular Whisper. Existing approaches either require addi…

eess.AS2024

Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution

Chin-Yun Yu, Sung-Lin Yeh, György Fazekas +1

Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given l…

cs.LG2024

Open-Source Conversational AI with SpeechBrain 1.0

Mirco Ravanelli, Titouan Parcollet, Adel Moumen +30

SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker re…