102 citations · 111 across the 2 of their papers we have counts for
5 papers · 1 filter
TMT: A Transformer-based Modal Translator for Improving Multimodal Sequence Representations in Audio Visual Scene-aware Dialog
Wubo Li, Dongwei Jiang, Wei Zou +1
Audio Visual Scene-aware Dialog (AVSD) is a task to generate responses when discussing about a given video. The previous state-of-the-art model shows superior performance for this…
Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
Dongwei Jiang, Wubo Li, Miao Cao +2
Self-supervised visual pretraining has shown significant progress recently. Among those methods, SimCLR greatly advanced the state of the art in self-supervised and semi-supervised…
Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
Dongwei Jiang, Xiaoning Lei, Wubo Li +4
Speech recognition technologies are gaining enormous popularity in various industrial applications. However, building a good speech recognition system usually requires large amount…
Towards End-to-End Code-Switching Speech Recognition
Ne Luo, Dongwei Jiang, Shuaijiang Zhao +3
Code-switching speech recognition has attracted an increasing interest recently, but the need for expert linguistic knowledge has always been a big issue. End-to-end automatic spee…
A comparable study of modeling units for end-to-end Mandarin speech recognition
Wei Zou, Dongwei Jiang, Shuaijiang Zhao +1
End-To-End speech recognition have become increasingly popular in mandarin speech recognition and achieved delightful performance. Mandarin is a tonal language which is different f…