activity
20152022
most citedHow Much Can CLIP Benefit Vision-and-Language Tasks?

153 citations · 601 across the 67 of their papers we have counts for

collaborators
Showing 2022Show all

14 papers · 1 filter

cs.CL2022

Mutual Exclusivity Training and Primitive Augmentation to Induce Compositionality

Yichen Jiang, Xiang Zhou, Mohit Bansal

Recent datasets expose the lack of the systematic generalization ability in standard sequence-to-sequence models. In this work, we analyze this behavior of seq2seq models and ident…

cs.CV20221 cited

Perceiver-VL: Efficient Vision-and-Language Modeling with Iterative Latent Attention

Zineng Tang, Jaemin Cho, Jie Lei +1

We present Perceiver-VL, a vision-and-language framework that efficiently handles high-dimensional multimodal inputs such as long videos and text. Powered by the iterative latent c…

cs.CL20221 cited

Are Hard Examples also Harder to Explain? A Study with Human and Model-Generated Explanations

Swarnadeep Saha, Peter Hase, Nazneen Rajani +1

Recent work on explainable NLP has shown that few-shot prompting can enable large pretrained language models (LLMs) to generate grammatical and factual natural language explanation…

cs.CL2022

Evaluating and Improving Factuality in Multimodal Abstractive Summarization

David Wan, Mohit Bansal

Current metrics for evaluating factuality for abstractive document summarization have achieved high correlations with human judgment, but they do not account for the vision modalit…

cs.CV202216 cited

TVLT: Textless Vision-Language Transformer

Zineng Tang, Jaemin Cho, Yixin Nie +1

In this work, we present the Textless Vision-Language Transformer (TVLT), where homogeneous transformer blocks take raw visual and audio inputs for vision-and-language representati…

cs.CV20225 cited

StoryDALL-E: Adapting Pretrained Text-to-Image Transformers for Story Continuation

Adyasha Maharana, Darryl Hannan, Mohit Bansal

Recent advances in text-to-image synthesis have led to large pretrained transformers with excellent capabilities to generate visualizations from a given text. However, these models…