119 citations · 524 across the 29 of their papers we have counts for
4 papers · 1 filter
Audio-Visual Scene-Aware Dialog and Reasoning using Audio-Visual Transformers with Joint Student-Teacher Learning
Ankit P. Shah, Shijie Geng, Peng Gao +5
In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at bo…
Scalable Transformers for Neural Machine Translation
Peng Gao, Shijie Geng, Yu Qiao +3
Transformer has been widely adopted in Neural Machine Translation (NMT) because of its large capacity and parallel training of sequence generation. However, the deployment of Trans…
RomeBERT: Robust Training of Multi-Exit BERT
Shijie Geng, Peng Gao, Zuohui Fu +1
BERT has achieved superior performances on Natural Language Understanding (NLU) tasks. However, BERT possesses a large number of parameters and demands certain resources to deploy.…
Multi-Pass Transformer for Machine Translation
Peng Gao, Chiori Hori, Shijie Geng +2
In contrast with previous approaches where information flows only towards deeper layers of a stack, we consider a multi-pass transformer (MPT) architecture in which earlier layers…