4 papers · 1 filter
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
Sung-Lin Yeh, Wei Zhou, Gil Keren +6
Recent speech language models rely on encoders that are optimized separately from autoregressive models. Since these encoders are unaware of the downstream objectives, the extracte…
Learning Speech Representations with Variational Predictive Coding
Sung-Lin Yeh, Peter Bell, Hao Tang
Despite being the best known objective for learning speech representations, the HuBERT objective has not been further developed and improved. We argue that it is the lack of an und…
Whisper Has an Internal Word Aligner
Sung-Lin Yeh, Yen Meng, Hao Tang
There is an increasing interest in obtaining accurate word-level timestamps from strong automatic speech recognizers, in particular Whisper. Existing approaches either require addi…
Estimating the Completeness of Discrete Speech Units
Sung-Lin Yeh, Hao Tang
Representing speech with discrete units has been widely used in speech codec and speech generation. However, there are several unverified claims about self-supervised discrete unit…