11 papers
The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models
Dehui Gao, Zhixian Zhao, Zhennan Lin +14
Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have created new opportunities for wearable speech interfaces, with smart glasses providing an egocentri…
SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription
Yuhang Dai, Haopeng Lin, Zhennan Lin +10
Recent advances in Automatic Speech Recognition (ASR) and Large Language Models (LLMs) have significantly improved speech understanding capabilities. However, multi-speaker speech…
Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition
Yuhang Dai, Haopeng Lin, Jiale Qian +9
Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limi…
SoulX-Duplug: Plug-and-Play Streaming State Prediction Module for Realtime Full-Duplex Speech Conversation
Ruiqi Yan, Wenxi Chen, Zhanxun Liu +17
Recent advances in spoken dialogue systems have brought increased attention to human-like full-duplex voice interactions. However, our comprehensive review of this field reveals se…
SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
Jiale Qian, Hao Meng, Tian Zheng +17
While recent years have witnessed rapid progress in speech synthesis, open-source singing voice synthesis (SVS) systems still face significant barriers to industrial deployment, pa…
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
Zhanxun Liu, Yifan Duan, Mengmeng Wang +15
We present X-Talk, an open-source framework that champions a decoupled, modular design for LLM-driven speech-to-speech (S2S) systems. While the dominant trend favors end-to-end (E2…