12 citations · 22 across the 5 of their papers we have counts for
5 papers
SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation
Yunrui Cai, Xu Li, Yucheng Zhou +8
Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coheren…
FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech
Shuoyi Zhou, Yixuan Zhou, Peiji Yang +4
Controllable text-to-speech (TTS) has become a key research focus. However, methods based on either reference speech or text descriptions lack flexibility and precise control, and…
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
Yuanyuan Wang, Dongchao Yang, Yiwen Shao +5
Extending pre-trained text Large Language Models (LLMs)'s speech understanding or generation abilities by introducing various effective speech tokens has attracted great attention…
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
Yixuan Zhou, Xiaoyu Qin, Zeyu Jin +5
Recent AIGC systems possess the capability to generate digital multimedia content based on human language instructions, such as text, image and video. However, when it comes to spe…
SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
Zeyu Jin, Jia Jia, Qixin Wang +5
Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elab…