activity
20242026
collaborators
Showing cs.SDShow all

13 papers · 1 filter

cs.SD2026

H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR

Yujie Guo, Jiaming Zhou, Yuhang Jia +2

Multi-talker Automatic Speech Recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech, particularly under complex high-overlap conditions. Wh…

cs.SD2026

EchoEdit: Stabilizing Inversion-Free Audio Editing via Optimal Transport Geometry

Zhongyuan Fu, Yuhang Jia, Hui Wang +6

Text-guided audio editing with pretrained generative models is commonly implemented through inversion or noising. This topology induces a structural trade-off, as stronger edits re…

cs.SD2026

CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

Junyang Chen, Yuhang Jia, Hui Wang +3

Speech editing and zero-shot Text-to-Speech (TTS) share a similar generative foundation conditioned on speech prompts, yet speech editing demands far stricter local acoustic consis…

cs.SD2026

DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding

Jiaming Zhou, Xuxin Cheng, Shiwan Zhao +5

Autoregressive (AR) large audio language models (LALMs) such as Qwen-2.5-Omni have achieved strong performance on audio understanding and interaction, but scaling them remains cost…

cs.SD2026

CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models

Junyang Chen, Yuhang Jia, Hui Wang +2

Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems rely on explicit temporal alignment and complex preprocessing.…

cs.SD2025

AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation

Hui Wang, Jinghua Zhao, Junyang Cheng +5

Text-to-audio (TTA) generation is advancing rapidly, but evaluation remains challenging because human listening studies are expensive and existing automatic metrics capture only li…