6 papers
SCDiar: a streaming diarization system based on speaker change detection and speech recognition
Naijun Zheng, Xucheng Wan, Kai Liu +1
In hours-long meeting scenarios, real-time speech stream often struggles with achieving accurate speaker diarization, commonly leading to speaker identification and speaker count e…
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
He Wang, Xucheng Wan, Naijun Zheng +4
Code-switching automatic speech recognition (ASR) aims to transcribe speech that contains two or more languages accurately. To better capture language-specific speech representatio…
XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
Xucheng Wan, Naijun Zheng, Kai Liu +1
Contextualized ASR models have been demonstrated to effectively improve the recognition accuracy of uncommon phrases when a predefined phrase list is available. However, these mode…
An efficient text augmentation approach for contextualized Mandarin speech recognition
Naijun Zheng, Xucheng Wan, Kai Liu +2
Although contextualized automatic speech recognition (ASR) systems are commonly used to improve the recognition of uncommon words, their effectiveness is hindered by the inherent l…
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
Bingshen Mu, Yangze Li, Qijie Shao +5
Despite notable advancements in automatic speech recognition (ASR), performance tends to degrade when faced with adverse conditions. Generative error correction (GER) leverages the…
Enhancing Lip Reading with Multi-Scale Video and Multi-Encoder
He Wang, Pengcheng Guo, Xucheng Wan +2
Automatic lip-reading (ALR) aims to automatically transcribe spoken content from a speaker's silent lip motion captured in video. Current mainstream lip-reading approaches only use…