activity
20242026
collaborators

6 papers

eess.AS2026

Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models

Jing-Xuan Zhang, Genshun Wan, Jin Li +3

While speech foundation models (SFMs) have demonstrated remarkable performance in audio-only tasks, their adaptation to multimodal scenarios remains underexplored. This work presen…

eess.AS2026

Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization

Genshun Wan, Wenhui Zhang, Jing-Xuan Zhang +3

Recent advances have demonstrated the potential of decoderonly large language models (LLMs) for automatic speech recognition (ASR). However, enabling streaming recognition within t…

eess.AS2025

Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models

Jing-Xuan Zhang, Genshun Wan, Jianqing Gao +1

Audio-visual representation learning is crucial for advancing multimodal speech processing tasks, such as lipreading and audio-visual speech recognition. Recently, speech foundatio…

cs.CL2024

Deep CLAS: Deep Contextual Listen, Attend and Spell

Mengzhi Wang, Shifu Xiong, Genshun Wan +3

Contextual-LAS (CLAS) has been shown effective in improving Automatic Speech Recognition (ASR) of rare words. It relies on phrase-level contextual modeling and attention-based rele…

cs.CL2024

Lightweight Transducer Based on Frame-Level Criterion

Genshun Wan, Mengzhi Wang, Tingzhi Mao +2

The transducer model trained based on sequence-level criterion requires a lot of memory due to the generation of the large probability matrix. We proposed a lightweight transducer…

eess.AS2024

The USTC-NERCSLIP Systems for the CHiME-8 NOTSOFAR-1 Challenge

Shutong Niu, Ruoyu Wang, Jun Du +17

This technical report outlines our submission system for the CHiME-8 NOTSOFAR-1 Challenge. The primary difficulty of this challenge is the dataset recorded across various conferenc…