4 papers
The USTC-NERCSLIP Systems for the CHiME-9 MCoRec Challenge
Ya Jiang, Ruoyu Wang, Jingxuan Zhang +19
This report details our submission to the CHiME-9 MCoRec Challenge on recognizing and clustering multiple concurrent natural conversations within indoor social settings. Unlike con…
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
Genshun Wan, Wenhui Zhang, Jing-Xuan Zhang +3
Recent advances have demonstrated the potential of decoderonly large language models (LLMs) for automatic speech recognition (ASR). However, enabling streaming recognition within t…
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
Jing-Xuan Zhang, Genshun Wan, Jin Li +3
While speech foundation models (SFMs) have demonstrated remarkable performance in audio-only tasks, their adaptation to multimodal scenarios remains underexplored. This work presen…
Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation
Jing-Xuan Zhang, Tingzhi Mao, Longjiang Guo +2
Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leve…