4 papers
Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis
Jinchuan Tian, Haoran Wang, Siddhant Arora +6
Classical TTS systems typically rely on rigid input formats and predefined metadata slots, limiting their ability to fulfill flexible user requirements. This paper introduces Bagpi…
Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
Jinchuan Tian, Haoran Wang, Bo-Hao Su +14
Current audio foundation models typically rely on rigid, task-specific supervision (e.g., speech recognition), addressing isolated factors of audio rather than the whole. In contra…
Online Predictive Coding for Dual-Mode Self-Supervised Speech Model
Keita Goto, Takashi Maekaku, Jin Sakuma +3
Dual-mode self-supervised speech models are pre-trained to handle streaming and non-streaming conditions simultaneously. However, their attention is computed over different context…
Online Register for Dual-Mode Self-Supervised Speech Models: Mitigating The Lack of Future Context
Keita Goto, Takashi Maekaku, Jin Sakuma +3
Dual-mode self-supervised speech models (S3Ms), which jointly pre-trained in the offline and online mode, suffer from attention mismatch in streaming scenarios due to missing futur…