most citedInteractive Audio-text Representation for Automated Audio Captioning with Contrastive Learning

9 citations · 12 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS20222 cited

Robust Data2vec: Noise-robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive Learning

Qiu-Shi Zhu, Long Zhou, Jie Zhang +3

Self-supervised pre-training methods based on contrastive learning or regression tasks can utilize more unlabeled data to improve the performance of automatic speech recognition (A…

cs.CL2022

Self-critical Sequence Training for Automatic Speech Recognition

Chen Chen, Yuchen Hu, Nana Hou +3

Although automatic speech recognition (ASR) task has gained remarkable success by sequence-to-sequence models, there are two main mismatches between its training and testing that m…

cs.SD20229 cited

Interactive Audio-text Representation for Automated Audio Captioning with Contrastive Learning

Chen Chen, Nana Hou, Yuchen Hu +3

Automated Audio captioning (AAC) is a cross-modal task that generates natural language to describe the content of input audio. Most prior works usually extract single-modality acou…

cs.SD2022

Noise-robust Speech Recognition with 10 Minutes Unparalleled In-domain Data

Chen Chen, Nana Hou, Yuchen Hu +2

Noise-robust speech recognition systems require large amounts of training data including noisy speech data and corresponding transcripts to achieve state-of-the-art performances in…

cs.CL20211 cited

The USTC-NELSLIP Systems for Simultaneous Speech Translation Task at IWSLT 2021

Dan Liu, Mengge Du, Xiaoxi Li +2

This paper describes USTC-NELSLIP's submissions to the IWSLT2021 Simultaneous Speech Translation task. We proposed a novel simultaneous translation model, Cross Attention Augmented…