18 papers · 1 filter
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
Junyang Chen, Yuhang Jia, Hui Wang +2
Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems rely on explicit temporal alignment and complex preprocessing.…
Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs
Yuhang Jia, Xu Zhang, Yang Chen +3
Automatic mean opinion score (MOS) prediction serves as a principled alternative to both subjective listening tests and objective metrics, providing scalable and consistent audio e…
EchoEdit: Stabilizing Inversion-Free Audio Editing via Optimal Transport Geometry
Zhongyuan Fu, Yuhang Jia, Hui Wang +6
Text-guided audio editing with pretrained generative models is commonly implemented through inversion or noising. This topology induces a structural trade-off, as stronger edits re…
FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision
Shiyao Wang, Xijuan Zeng, Hui Wang +4
We present FoleyGenEx, a unified video-to-audio (VTA) framework integrating multi-modal control, frame-level temporal alignment, and fine-grained semantics, enabling synchronized,…
CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS
Junyang Chen, Yuhang Jia, Hui Wang +3
Speech editing and zero-shot Text-to-Speech (TTS) share a similar generative foundation conditioned on speech prompts, yet speech editing demands far stricter local acoustic consis…
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
Hui Wang, Jinghua Zhao, Yifan Yang +9
Generative speech technologies are progressing rapidly, but evaluating the perceptual quality of synthetic speech remains a core challenge. Existing methods typically rely on scala…