3 citations · 4 across the 5 of their papers we have counts for
5 papers
X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System
Yuxiang Zhao, Yichi Zhang, Yanjie An +10
Real-time speech-to-speech translation (S2ST) systems must balance translation quality, latency, speech naturalness, and speaker consistency. Publicly documented S2ST systems have…
The SJTU X-LANCE Lab System for MSR Challenge 2025
Jinxuan Zhu, Hao Qiu, Haina Zhu +3
This report describes the system submitted to the music source restoration (MSR) Challenge 2025. Our approach is composed of sequential BS-RoFormers, each dealing with a single tas…
Audio ControlNet for Fine-Grained Audio Generation and Editing
Haina Zhu, Yao Xiao, Xiquan Li +5
We study the fine-grained text-to-audio (T2A) generation task. While recent models can synthesize high-quality audio from text descriptions, they often lack precise control over at…
SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
Ziyang Ma, Guanrou Yang, Wenxi Chen +19
The recent surge in open-source Multimodal Large Language Models (MLLM) frameworks, such as LLaVA, provides a convenient kickoff for artificial intelligence developers and research…
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
Zhanxun Liu, Yifan Duan, Mengmeng Wang +15
We present X-Talk, an open-source framework that champions a decoupled, modular design for LLM-driven speech-to-speech (S2S) systems. While the dominant trend favors end-to-end (E2…