Publications (23)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
Jiawen Kang, Lingwei Meng, Mingyu Cui +4
End-to-end multi-talker speech recognition has garnered great interest as an effective approach to directly transcribe overlapped speech from multiple speakers. Current methods typ…
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
Dongchao Yang, Haohan Guo, Yuanyuan Wang +5
The Large Language models (LLMs) have demonstrated supreme capabilities in text understanding and generation, but cannot be directly applied to cross-modal tasks without fine-tunin…
A New GAN-based End-to-End TTS Training Algorithm
Haohan Guo, Frank K. Soong, Lei He +1
End-to-end, autoregressive model-based TTS has shown significant performance improvements over the conventional one. However, the autoregressive module training is affected by the…
Exploiting Syntactic Features in a Parsed Tree to Improve End-to-End TTS
Haohan Guo, Frank K. Soong, Lei He +1
The end-to-end TTS, which can predict speech directly from a given sequence of graphemes or phonemes, has shown improved performance over the conventional TTS. However, its predict…
Improving Adversarial Waveform Generation based Singing Voice Conversion with Harmonic Signals
Haohan Guo, Zhiping Zhou, Fanbo Meng +1
Adversarial waveform generation has been a popular approach as the backend of singing voice conversion (SVC) to generate high-quality singing audio. However, the instability of GAN…
HeartMuLa: A Family of Open Sourced Music Foundation Models
Dongchao Yang, Yuxin Xie, Yuguo Yin +26
We present a family of open-source Music Foundation Models designed to advance large-scale music understanding and generation across diverse tasks and modalities. Our framework con…