5 papers
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
Sung-Lin Yeh, Wei Zhou, Gil Keren +6
Recent speech language models rely on encoders that are optimized separately from autoregressive models. Since these encoders are unaware of the downstream objectives, the extracte…
Learning Speech Representations with Variational Predictive Coding
Sung-Lin Yeh, Peter Bell, Hao Tang
Despite being the best known objective for learning speech representations, the HuBERT objective has not been further developed and improved. We argue that it is the lack of an und…
Whisper Has an Internal Word Aligner
Sung-Lin Yeh, Yen Meng, Hao Tang
There is an increasing interest in obtaining accurate word-level timestamps from strong automatic speech recognizers, in particular Whisper. Existing approaches either require addi…
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
Chin-Yun Yu, Sung-Lin Yeh, György Fazekas +1
Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given l…
Open-Source Conversational AI with SpeechBrain 1.0
Mirco Ravanelli, Titouan Parcollet, Adel Moumen +30
SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker re…