activity
20182022
most citedX-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers

24 citations · 44 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV20221 cited

Perceiver-VL: Efficient Vision-and-Language Modeling with Iterative Latent Attention

Zineng Tang, Jaemin Cho, Jie Lei +1

We present Perceiver-VL, a vision-and-language framework that efficiently handles high-dimensional multimodal inputs such as long videos and text. Powered by the iterative latent c…

cs.CV202216 cited

TVLT: Textless Vision-Language Transformer

Zineng Tang, Jaemin Cho, Yixin Nie +1

In this work, we present the Textless Vision-Language Transformer (TVLT), where homogeneous transformer blocks take raw visual and audio inputs for vision-and-language representati…

cs.CL20213 cited

VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer

Zineng Tang, Jaemin Cho, Hao Tan +1

Since visual perception can give rich information beyond text descriptions for world understanding, there has been increasing interest in leveraging visual grounding for language l…

cs.CL2021

Unifying Vision-and-Language Tasks via Text Generation

Jaemin Cho, Jie Lei, Hao Tan +1

Existing methods for vision-and-language learning typically require designing task-specific architectures and objectives for each task. For example, a multi-label answer classifier…

cs.CV202024 cited

X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers

Jaemin Cho, Jiasen Lu, Dustin Schwenk +2

Mirroring the success of masked language models, vision-and-language counterparts like ViLBERT, LXMERT and UNITER have achieved state of the art performance on a variety of multimo…

cs.CL2019

Mixture Content Selection for Diverse Sequence Generation

Jaemin Cho, Minjoon Seo, Hannaneh Hajishirzi

Generating diverse sequences is important in many NLP applications such as question generation or summarization that exhibit semantically one-to-many relationships between source a…