activity
20182022
most citedImproving Transformer-based Speech Recognition Using Unsupervised Pre-training

102 citations · 111 across the 7 of their papers we have counts for

collaborators

15 papers

eess.AS2022

Audio Deep Fake Detection System with Neural Stitching for ADD 2022

Rui Yan, Cheng Wen, Shuran Zhou +3

This paper describes our best system and methodology for ADD 2022: The First Audio Deep Synthesis Detection Challenge\cite{Yi2022ADD}. The very same system was used for both two ro…

eess.AS2022

Time Domain Adversarial Voice Conversion for ADD 2022

Cheng Wen, Tingwei Guo, Xingjun Tan +5

In this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) sy…

cs.SD2022

Audio-Visual Wake Word Spotting System For MISP Challenge 2021

Yanguang Xu, Jianwei Sun, Yang Han +7

This paper presents the details of our system designed for the Task 1 of Multimodal Information Based Speech Processing (MISP) Challenge 2021. The purpose of Task 1 is to leverage…

eess.AS2021

Semantic Data Augmentation for End-to-End Mandarin Speech Recognition

Jianwei Sun, Zhiyuan Tang, Hengxin Yin +6

End-to-end models have gradually become the preferred option for automatic speech recognition (ASR) applications. During the training of end-to-end ASR, data augmentation is a quit…

cs.CL2020

TMT: A Transformer-based Modal Translator for Improving Multimodal Sequence Representations in Audio Visual Scene-aware Dialog

Wubo Li, Dongwei Jiang, Wei Zou +1

Audio Visual Scene-aware Dialog (AVSD) is a task to generate responses when discussing about a given video. The previous state-of-the-art model shows superior performance for this…

cs.CL2020

Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning

Dongwei Jiang, Wubo Li, Miao Cao +2

Self-supervised visual pretraining has shown significant progress recently. Among those methods, SimCLR greatly advanced the state of the art in self-supervised and semi-supervised…