5 papers
SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
Jiale Qian, Hao Meng, Tian Zheng +17
While recent years have witnessed rapid progress in speech synthesis, open-source singing voice synthesis (SVS) systems still face significant barriers to industrial deployment, pa…
BrainRVQ: A High-Fidelity EEG Foundation Model via Dual-Domain Residual Quantization and Hierarchical Autoregression
Mingzhe Cui, Tao Chen, Yang Jiao +4
Developing foundation models for electroencephalography (EEG) remains challenging due to the signal's low signal-to-noise ratio and complex spectro-temporal non-stationarity. Exist…
LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
Bingshen Mu, Xian Shi, Xiong Wang +3
Forced alignment (FA) predicts start and end timestamps for words or characters in speech, but existing methods are language-specific and prone to cumulative temporal shifts. The m…
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
Wenjie Tian, Bingshen Mu, Guobin Ma +3
Automatic speech recognition (ASR) systems based on large language models (LLMs) achieve superior performance by leveraging pretrained LLMs as decoders, but their token-by-token ge…
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
Zhanxun Liu, Yifan Duan, Mengmeng Wang +15
We present X-Talk, an open-source framework that champions a decoupled, modular design for LLM-driven speech-to-speech (S2S) systems. While the dominant trend favors end-to-end (E2…