1.6k citations
- Ministry of Industry and Information TechnologyCN16 papers
- Tsinghua UniversityCN16 papers
- Chinese Academy of SciencesCN15 papers
- Xi'an Jiaotong UniversityCN15 papers
- Xidian UniversityCN12 papers
- Australian National UniversityAU11 papers
- Moscow Institute of Physics and TechnologyRU11 papers
- University of Chinese Academy of SciencesCN10 papers
- Nankai UniversityCN9 papers
- Centre National de la Recherche ScientifiqueFR7 papers
- Fudan UniversityCN7 papers
- Instituto de Ciencia de Materiales de MadridES7 papers
16 papers · 1 filter
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition
He Wang, Pengcheng Guo, Pan Zhou +1
While automatic speech recognition (ASR) systems degrade significantly in noisy environments, audio-visual speech recognition (AVSR) systems aim to complement the audio stream with…
Decoupling and Interacting Multi-Task Learning Network for Joint Speech and Accent Recognition
Qijie Shao, Pengcheng Guo, Jinghao Yan +2
Accents, as variations from standard pronunciation, pose significant challenges for speech recognition systems. Although joint automatic speech recognition (ASR) and accent recogni…
Multi-Task Deep Residual Echo Suppression with Echo-aware Loss
Shimin Zhang, Ziteng Wang, Jiayao Sun +4
This paper introduces the NWPU Team's entry to the ICASSP 2022 AEC Challenge. We take a hybrid approach that cascades a linear AEC with a neural post-filter. The former is used to…
Conformer-based End-to-end Speech Recognition With Rotary Position Embedding
Shengqiang Li, Menglong Xu, Xiao-Lei Zhang
Transformer-based end-to-end speech recognition models have received considerable attention in recent years due to their high training speed and ability to model a long-range globa…
Audio Description from Image by Modal Translation Network
Hailong Ning, Xiangtao Zheng, Yuan Yuan +1
Audio is the main form for the visually impaired to obtain information. In reality, all kinds of visual data always exist, but audio data does not exist in many cases. In order to…
An Asynchronous WFST-Based Decoder For Automatic Speech Recognition
Hang Lv, Zhehuai Chen, Hainan Xu +3
We introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech…