activity
20242026
collaborators
Showing cs.SDShow all

18 papers · 1 filter

cs.SD2026

CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models

Junyang Chen, Yuhang Jia, Hui Wang +2

Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems rely on explicit temporal alignment and complex preprocessing.…

cs.SD2026

Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs

Yuhang Jia, Xu Zhang, Yang Chen +3

Automatic mean opinion score (MOS) prediction serves as a principled alternative to both subjective listening tests and objective metrics, providing scalable and consistent audio e…

cs.SD2026

EchoEdit: Stabilizing Inversion-Free Audio Editing via Optimal Transport Geometry

Zhongyuan Fu, Yuhang Jia, Hui Wang +6

Text-guided audio editing with pretrained generative models is commonly implemented through inversion or noising. This topology induces a structural trade-off, as stronger edits re…

cs.SD2026

FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision

Shiyao Wang, Xijuan Zeng, Hui Wang +4

We present FoleyGenEx, a unified video-to-audio (VTA) framework integrating multi-modal control, frame-level temporal alignment, and fine-grained semantics, enabling synchronized,…

cs.SD2026

CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

Junyang Chen, Yuhang Jia, Hui Wang +3

Speech editing and zero-shot Text-to-Speech (TTS) share a similar generative foundation conditioned on speech prompts, yet speech editing demands far stricter local acoustic consis…

cs.SD2026

SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation

Hui Wang, Jinghua Zhao, Yifan Yang +9

Generative speech technologies are progressing rapidly, but evaluating the perceptual quality of synthetic speech remains a core challenge. Existing methods typically rely on scala…