37 citations · 38 across the 3 of their papers we have counts for
3 papers
cs.LG2022★ 37 cited
Mega: Moving Average Equipped Gated Attention
Xuezhe Ma, Chunting Zhou, Xiang Kong +5
The design choices in the Transformer attention mechanism, including weak inductive bias and quadratic computational complexity, have limited its application for modeling long sequ…
cs.CV2022★ 1 cited
Training Vision-Language Transformers from Captions
Liangke Gui, Yingshan Chang, Qiuyuan Huang +4
Vision-Language Transformers can be learned without low-level human labels (e.g. class labels, bounding boxes, etc). Existing work, whether explicitly utilizing bounding boxes or p…
cs.CL2021
KAT: A Knowledge Augmented Transformer for Vision-and-Language
Liangke Gui, Borui Wang, Qiuyuan Huang +3
The primary focus of recent work with largescale transformers has been on optimizing the amount of information packed into the model's parameters. In this work, we ask a different…