6 papers
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
Jing-Xuan Zhang, Genshun Wan, Jin Li +3
While speech foundation models (SFMs) have demonstrated remarkable performance in audio-only tasks, their adaptation to multimodal scenarios remains underexplored. This work presen…
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
Genshun Wan, Wenhui Zhang, Jing-Xuan Zhang +3
Recent advances have demonstrated the potential of decoderonly large language models (LLMs) for automatic speech recognition (ASR). However, enabling streaming recognition within t…
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
Jing-Xuan Zhang, Genshun Wan, Jianqing Gao +1
Audio-visual representation learning is crucial for advancing multimodal speech processing tasks, such as lipreading and audio-visual speech recognition. Recently, speech foundatio…
Deep CLAS: Deep Contextual Listen, Attend and Spell
Mengzhi Wang, Shifu Xiong, Genshun Wan +3
Contextual-LAS (CLAS) has been shown effective in improving Automatic Speech Recognition (ASR) of rare words. It relies on phrase-level contextual modeling and attention-based rele…
Lightweight Transducer Based on Frame-Level Criterion
Genshun Wan, Mengzhi Wang, Tingzhi Mao +2
The transducer model trained based on sequence-level criterion requires a lot of memory due to the generation of the large probability matrix. We proposed a lightweight transducer…
The USTC-NERCSLIP Systems for the CHiME-8 NOTSOFAR-1 Challenge
Shutong Niu, Ruoyu Wang, Jun Du +17
This technical report outlines our submission system for the CHiME-8 NOTSOFAR-1 Challenge. The primary difficulty of this challenge is the dataset recorded across various conferenc…