acoustic features 1full-duplex dialogue 1low latency 1noise robustness 1streaming ctc 1turn detection 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.SD2026
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection
Chengyou Wang, Hongfei Xue, Mingchen Shao +9
FastTurn is a unified framework that combines streaming CTC decoding with acoustic and semantic cues to achieve low-latency, robust turn detection for real-time full‑duplex spoken…
eess.AS2026
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Zhaokai Sun, Shuai Wang, Zhennan Lin +6
Spoken Language Understanding (SLU) is moving from task-specific pipelines toward large audio language models (LALMs) that generate natural-language responses. However, existing sp…
eess.AS2026
UrduSpeech: A 156-Hour Urdu Speech Corpus with 12-Dimension Paralinguistic Annotations
Attia Nafees ul Haq, Zeyu Zhu, Jingbin Hu +2
Despite 230 million speakers, Urdu remains critically under-resourced in speech technology. We introduce UrduSpeech: a large high-fidelity Urdu corpus comprising 156 hours of audio…