activity
20242026
most citedMMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

1 citations · 1 across the 6 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2026

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

Yunrui Cai, Xu Li, Yucheng Zhou +8

Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coheren…

cs.SD2026

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

Dingdong Wang, Shujie Liu, Tianhua Zhang +3

Emotional information in speech plays a unique role in multimodal perception. However, current Speech Large Language Models (SpeechLLMs), similar to conventional speech emotion rec…

cs.SD2026

V2A-DPO: Omni-Preference Optimization for Video-to-Audio Generation

Nolan Chan, Timmy Gang, Yongqian Wang +2

This paper introduces V2A-DPO, a novel Direct Preference Optimization (DPO) framework tailored for flow-based video-to-audio generation (V2A) models, incorporating key adaptations…

cs.SD2025

InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training

Dingdong Wang, Jin Xu, Ruihang Chu +6

Recent advancements in speech large language models (SpeechLLMs) have attracted considerable attention. Nonetheless, current methods exhibit suboptimal performance in adhering to s…

cs.SD2024

CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction

Xueyuan Chen, Dongchao Yang, Dingdong Wang +3

Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech. It still suffers from low speaker similarity and poor prosody naturalness. In this pa…

cs.SD2024

SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Dongchao Yang, Dingdong Wang, Haohan Guo +3

In this study, we propose a simple and efficient Non-Autoregressive (NAR) text-to-speech (TTS) system based on diffusion, named SimpleSpeech. Its simpleness shows in three aspects:…