102 citations · 111 across the 3 of their papers we have counts for
8 papers
TMT: A Transformer-based Modal Translator for Improving Multimodal Sequence Representations in Audio Visual Scene-aware Dialog
Wubo Li, Dongwei Jiang, Wei Zou +1
Audio Visual Scene-aware Dialog (AVSD) is a task to generate responses when discussing about a given video. The previous state-of-the-art model shows superior performance for this…
Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
Dongwei Jiang, Wubo Li, Miao Cao +2
Self-supervised visual pretraining has shown significant progress recently. Among those methods, SimCLR greatly advanced the state of the art in self-supervised and semi-supervised…
DiDiSpeech: A Large Scale Mandarin Speech Corpus
Tingwei Guo, Cheng Wen, Dongwei Jiang +8
This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the…
Transformer based unsupervised pre-training for acoustic representation learning
Ruixiong Zhang, Haiwei Wu, Wubo Li +3
Recently, a variety of acoustic tasks and related applications arised. For many acoustic tasks, the labeled data size may be limited. To handle this problem, we propose an unsuperv…
A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition
Dongwei Jiang, Wubo Li, Ruixiong Zhang +5
Building a good speech recognition system usually requires large amounts of transcribed data, which is expensive to collect. To tackle this problem, many unsupervised pre-training…
Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
Dongwei Jiang, Xiaoning Lei, Wubo Li +4
Speech recognition technologies are gaining enormous popularity in various industrial applications. However, building a good speech recognition system usually requires large amount…