activity
20182022
most citedFocal Self-attention for Local-Global Interactions in Vision Transformers

268 citations · 465 across the 26 of their papers we have counts for

collaborators

18 papers

cs.CL20216 cited

CLUES: Few-Shot Learning Evaluation in Natural Language Understanding

Subhabrata Mukherjee, Xiaodong Liu, Guoqing Zheng +6

Most recent progress in natural language understanding (NLU) has been driven, in part, by benchmarks such as GLUE, SuperGLUE, SQuAD, etc. In fact, many NLU models have now matched…

cs.CL20212 cited

SYNERGY: Building Task Bots at Scale Using Symbolic Knowledge and Machine Teaching

Baolin Peng, Chunyuan Li, Zhu Zhang +3

In this paper we explore the use of symbolic knowledge and machine teaching to reduce human data labeling efforts in building neural task bots. We propose SYNERGY, a hybrid learnin…

cs.CV20214 cited

TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment

Jianwei Yang, Yonatan Bisk, Jianfeng Gao

Contrastive learning has been widely used to train transformer-based vision-language models for video-text alignment and multi-modal representation learning. This paper presents a…

cs.CL20211 cited

EmailSum: Abstractive Email Thread Summarization

Shiyue Zhang, Asli Celikyilmaz, Jianfeng Gao +1

Recent years have brought about an interest in the challenging task of summarizing conversation threads (meetings, online discussions, etc.). Such summaries help analysis of the lo…

cs.CV202117 cited

Image Scene Graph Generation (SGG) Benchmark

Xiaotian Han, Jianwei Yang, Houdong Hu +3

There is a surge of interest in image scene graph generation (object, attribute and relationship detection) due to the need of building fine-grained image understanding models that…

cs.CV2021268 cited

Focal Self-attention for Local-Global Interactions in Vision Transformers

Jianwei Yang, Chunyuan Li, Pengchuan Zhang +4

Recently, Vision Transformer and its variants have shown great promise on various computer vision tasks. The ability of capturing short- and long-range visual dependencies through…