102 citations · 111 across the 3 of their papers we have counts for
4 papers · 1 filter
TMT: A Transformer-based Modal Translator for Improving Multimodal Sequence Representations in Audio Visual Scene-aware Dialog
Wubo Li, Dongwei Jiang, Wei Zou +1
Audio Visual Scene-aware Dialog (AVSD) is a task to generate responses when discussing about a given video. The previous state-of-the-art model shows superior performance for this…
Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
Dongwei Jiang, Wubo Li, Miao Cao +2
Self-supervised visual pretraining has shown significant progress recently. Among those methods, SimCLR greatly advanced the state of the art in self-supervised and semi-supervised…
Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
Dongwei Jiang, Xiaoning Lei, Wubo Li +4
Speech recognition technologies are gaining enormous popularity in various industrial applications. However, building a good speech recognition system usually requires large amount…
A Multi-Modal Chinese Poetry Generation Model
Dayiheng Liu, Quan Guo, Wubo Li +1
Recent studies in sequence-to-sequence learning demonstrate that RNN encoder-decoder structure can successfully generate Chinese poetry. However, existing methods can only generate…