6 papers
X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System
Yuxiang Zhao, Yichi Zhang, Yanjie An +10
Real-time speech-to-speech translation (S2ST) systems must balance translation quality, latency, speech naturalness, and speaker consistency. Publicly documented S2ST systems have…
OpenSTBench: Beyond Semantic Evaluation for Speech Translation
Yanjie An, Yuxiang Zhao, Yichi Zhang +5
Speech translation systems increasingly span speech-to-text translation (S2TT), speech-to-speech translation (S2ST), offline translation, and streaming generation, producing output…
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
Qixi Zheng, Yuxiang Zhao, Tianrui Wang +7
Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content. Although recent systems have…
Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
Yunchong Xiao, Yuxiang Zhao, Ziyang Ma +4
The growing reliance on large-scale speech data has made privacy protection a critical concern. However, existing anonymization approaches often degrade data utility, for example b…
SoulX-Duplug: Plug-and-Play Streaming State Prediction Module for Realtime Full-Duplex Speech Conversation
Ruiqi Yan, Wenxi Chen, Zhanxun Liu +17
Recent advances in spoken dialogue systems have brought increased attention to human-like full-duplex voice interactions. However, our comprehensive review of this field reveals se…
Traceable TTS: Toward Watermark-Free TTS with Strong Traceability
Yuxiang Zhao, Yunchong Xiao, Yushen Chen +4
Recent advances in Text-To-Speech (TTS) technology have enabled synthetic speech to mimic human voices with remarkable realism, raising significant security concerns. This undersco…