13 papers
A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation
Liuyi Wang, Kai Sheng, Zongtao He +8
Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prior assumptions, limiting their robustness…
H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR
Yujie Guo, Jiaming Zhou, Yuhang Jia +2
Multi-talker Automatic Speech Recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech, particularly under complex high-overlap conditions. Wh…
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
Junyang Chen, Yuhang Jia, Hui Wang +2
Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems rely on explicit temporal alignment and complex preprocessing.…
Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs
Yuhang Jia, Xu Zhang, Yang Chen +3
Automatic mean opinion score (MOS) prediction serves as a principled alternative to both subjective listening tests and objective metrics, providing scalable and consistent audio e…
CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS
Junyang Chen, Yuhang Jia, Hui Wang +3
Speech editing and zero-shot Text-to-Speech (TTS) share a similar generative foundation conditioned on speech prompts, yet speech editing demands far stricter local acoustic consis…
Position: Towards Responsible Evaluation for Text-to-Speech
Yifan Yang, Hui Wang, Bing Han +4
Recent advances in text-to-speech (TTS) technology have enabled systems to generate speech that is often indistinguishable from human speech, bringing benefits to accessibility, co…