output
20062024
most citedStatistical physics of human cooperation

1.6k citations

Showing cs.SDShow all

16 papers · 1 filter

cs.SD202424 cited

MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition

He Wang, Pengcheng Guo, Pan Zhou +1

While automatic speech recognition (ASR) systems degrade significantly in noisy environments, audio-visual speech recognition (AVSR) systems aim to complement the audio stream with…

cs.SD202315 cited

Decoupling and Interacting Multi-Task Learning Network for Joint Speech and Accent Recognition

Qijie Shao, Pengcheng Guo, Jinghao Yan +2

Accents, as variations from standard pronunciation, pose significant challenges for speech recognition systems. Although joint automatic speech recognition (ASR) and accent recogni…

cs.SD202229 cited

Multi-Task Deep Residual Echo Suppression with Echo-aware Loss

Shimin Zhang, Ziteng Wang, Jiayao Sun +4

This paper introduces the NWPU Team's entry to the ICASSP 2022 AEC Challenge. We take a hybrid approach that cascades a linear AEC with a neural post-filter. The former is used to…

cs.SD2021

Conformer-based End-to-end Speech Recognition With Rotary Position Embedding

Shengqiang Li, Menglong Xu, Xiao-Lei Zhang

Transformer-based end-to-end speech recognition models have received considerable attention in recent years due to their high training speed and ability to model a long-range globa…

cs.SD202124 cited

Audio Description from Image by Modal Translation Network

Hailong Ning, Xiangtao Zheng, Yuan Yuan +1

Audio is the main form for the visually impaired to obtain information. In reality, all kinds of visual data always exist, but audio data does not exist in many cases. In order to…

cs.SD2021

An Asynchronous WFST-Based Decoder For Automatic Speech Recognition

Hang Lv, Zhehuai Chen, Hainan Xu +3

We introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech…