activity
20242026
most citedSpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description

12 citations · 22 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD2026

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

Yunrui Cai, Xu Li, Yucheng Zhou +8

Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coheren…

cs.SD2026

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech

Shuoyi Zhou, Yixuan Zhou, Peiji Yang +4

Controllable text-to-speech (TTS) has become a key research focus. However, methods based on either reference speech or text descriptions lack flexibility and precise control, and…

cs.SD2025

DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models

Yuanyuan Wang, Dongchao Yang, Yiwen Shao +5

Extending pre-trained text Large Language Models (LLMs)'s speech understanding or generation abilities by introducing various effective speech tokens has attracted great attention…

cs.SD202410 cited

VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling

Yixuan Zhou, Xiaoyu Qin, Zeyu Jin +5

Recent AIGC systems possess the capability to generate digital multimedia content based on human language instructions, such as text, image and video. However, when it comes to spe…

cs.MM202412 cited

SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description

Zeyu Jin, Jia Jia, Qixin Wang +5

Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elab…