collaborators

13 papers

cs.RO2026

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation

Liuyi Wang, Kai Sheng, Zongtao He +8

Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prior assumptions, limiting their robustness…

cs.SD2026

H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR

Yujie Guo, Jiaming Zhou, Yuhang Jia +2

Multi-talker Automatic Speech Recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech, particularly under complex high-overlap conditions. Wh…

cs.SD2026

CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models

Junyang Chen, Yuhang Jia, Hui Wang +2

Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems rely on explicit temporal alignment and complex preprocessing.…

cs.SD2026

Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs

Yuhang Jia, Xu Zhang, Yang Chen +3

Automatic mean opinion score (MOS) prediction serves as a principled alternative to both subjective listening tests and objective metrics, providing scalable and consistent audio e…

cs.SD2026

CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

Junyang Chen, Yuhang Jia, Hui Wang +3

Speech editing and zero-shot Text-to-Speech (TTS) share a similar generative foundation conditioned on speech prompts, yet speech editing demands far stricter local acoustic consis…

eess.AS2026

Position: Towards Responsible Evaluation for Text-to-Speech

Yifan Yang, Hui Wang, Bing Han +4

Recent advances in text-to-speech (TTS) technology have enabled systems to generate speech that is often indistinguishable from human speech, bringing benefits to accessibility, co…