activity
20152022
most citedHow Much Can CLIP Benefit Vision-and-Language Tasks?

153 citations · 601 across the 67 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV20221 cited

Perceiver-VL: Efficient Vision-and-Language Modeling with Iterative Latent Attention

Zineng Tang, Jaemin Cho, Jie Lei +1

We present Perceiver-VL, a vision-and-language framework that efficiently handles high-dimensional multimodal inputs such as long videos and text. Powered by the iterative latent c…

cs.CV202216 cited

TVLT: Textless Vision-Language Transformer

Zineng Tang, Jaemin Cho, Yixin Nie +1

In this work, we present the Textless Vision-Language Transformer (TVLT), where homogeneous transformer blocks take raw visual and audio inputs for vision-and-language representati…

cs.CV20225 cited

StoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story Continuation

Adyasha Maharana, Darryl Hannan, Mohit Bansal

Recent advances in text-to-image synthesis have led to large pretrained transformers with excellent capabilities to generate visualizations from a given text. However, these models…

cs.CV20225 cited

EnvEdit: Environment Editing for Vision-and-Language Navigation

Jialu Li, Hao Tan, Mohit Bansal

In Vision-and-Language Navigation (VLN), an agent needs to navigate through the environment based on natural language instructions. Due to limited available data for agent training…

cs.CV20226 cited

LoopITR: Combining Dual and Cross Encoder Architectures for Image-Text Retrieval

Jie Lei, Xinlei Chen, Ning Zhang +4

Dual encoders and cross encoders have been widely used for image-text retrieval. Between the two, the dual encoder encodes the image and text independently followed by a dot produc…

cs.CV2021153 cited

How Much Can CLIP Benefit Vision-and-Language Tasks?

Sheng Shen, Liunian Harold Li, Hao Tan +5

Most existing Vision-and-Language (V&L) models rely on pre-trained visual encoders, using a relatively small set of manually-annotated data (as compared to web-crawled data), to pe…