From the 1 of 20 linked papers with an AI index.
20 papers
On-Policy Self-Distillation for Multi-Dialect ASR: Mastering Dialects, Retaining Mandarin
Shuiyuan Wang, Bingshen Mu, Pengshen Zhang +6
Recent large-scale ASR models already achieve strong Mandarin recognition accuracy and have some ability to recognize Chinese dialects. However, their dialect recognition accuracy…
SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation
Hanke Xie, Haopeng Lin, Jiale Qian +13
Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic in…
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection
Chengyou Wang, Hongfei Xue, Mingchen Shao +9
FastTurn is a unified framework that combines streaming CTC decoding with acoustic and semantic cues to achieve low-latency, robust turn detection for real-time full‑duplex spoken…
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Zhaokai Sun, Shuai Wang, Zhennan Lin +6
Spoken Language Understanding (SLU) is moving from task-specific pipelines toward large audio language models (LALMs) that generate natural-language responses. However, existing sp…
Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge
Chengyou Wang, Hongfei Xue, Guojian Li +6
Full-duplex interaction, where speakers and listeners converse simultaneously, is a key element of human communication often missing from traditional spoken dialogue systems. These…
HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
Shuiyuan Wang, Zhixian Zhao, Hongfei Xue +5
Evaluating the emotional intelligence (EI) of audio language models (ALMs) is critical. However, existing benchmarks mostly rely on synthesized speech, are limited to single-turn i…