papers

Publications (23)

cs.SD2024

Cross-Speaker Encoding Network for Multi-Talker Speech Recognition

Jiawen Kang, Lingwei Meng, Mingyu Cui +4

End-to-end multi-talker speech recognition has garnered great interest as an effective approach to directly transcribe overlapped speech from multiple speakers. Current methods typ…

cs.SD2024

UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner

Dongchao Yang, Haohan Guo, Yuanyuan Wang +5

The Large Language models (LLMs) have demonstrated supreme capabilities in text understanding and generation, but cannot be directly applied to cross-modal tasks without fine-tunin…

cs.CL2019

A New GAN-based End-to-End TTS Training Algorithm

Haohan Guo, Frank K. Soong, Lei He +1

End-to-end, autoregressive model-based TTS has shown significant performance improvements over the conventional one. However, the autoregressive module training is affected by the…

cs.CL2019

Exploiting Syntactic Features in a Parsed Tree to Improve End-to-End TTS

Haohan Guo, Frank K. Soong, Lei He +1

The end-to-end TTS, which can predict speech directly from a given sequence of graphemes or phonemes, has shown improved performance over the conventional TTS. However, its predict…

cs.SD2022

Improving Adversarial Waveform Generation based Singing Voice Conversion with Harmonic Signals

Haohan Guo, Zhiping Zhou, Fanbo Meng +1

Adversarial waveform generation has been a popular approach as the backend of singing voice conversion (SVC) to generate high-quality singing audio. However, the instability of GAN…

cs.SD2026

HeartMuLa: A Family of Open Sourced Music Foundation Models

Dongchao Yang, Yuxin Xie, Yuguo Yin +26

We present a family of open-source Music Foundation Models designed to advance large-scale music understanding and generation across diverse tasks and modalities. Our framework con…