9 papers
Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future Context
Keita Goto, Takashi Maekaku, Jin Sakuma +3
Dual-mode self-supervised speech models (S3Ms), which jointly pre-trained in the offline and online mode, suffer from attention mismatch in streaming scenarios due to missing futur…
Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition
Bo-Hao Su, Hui-Ying Shih, Jinchuan Tian +4
Speech Emotion Recognition (SER) is typically trained and evaluated on majority-voted labels, which simplifies benchmarking but masks subjectivity and provides little transparency…
Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
Shih-Heng Wang, Jiatong Shi, Jinchuan Tian +2
This paper investigates three crucial yet underexplored aspects of the generalization capabilities of neural audio codecs (NACs): (i) whether NACs can generalize to unseen language…
PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
Jiatong Shi, Haoran Wang, William Chen +4
Neural speech codecs have achieved strong performance in low-bitrate compression, but residual vector quantization (RVQ) often suffers from unstable training and ineffective decomp…
BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
Haoran Wang, Jiatong Shi, Jinchuan Tian +3
Neural audio codecs have recently enabled high-fidelity reconstruction at high compression rates, especially for speech. However, speech and non-speech audio exhibit fundamentally…
Evaluating Self-Supervised Speech Models via Text-Based LLMS
Takashi Maekaku, Keita Goto, Jinchuan Tian +2
Self-Supervised Learning (SSL) has gained traction for its ability to learn rich representations with low labeling costs, applicable across diverse downstream tasks. However, asses…