activity
20152020
most citedVisual Madlibs: Fill in the blank Image Generation and Question Answering

80 citations · 162 across the 6 of their papers we have counts for

collaborators

13 papers

cs.CL2020

What is More Likely to Happen Next? Video-and-Language Future Event Prediction

Jie Lei, Licheng Yu, Tamara L. Berg +1

Given a video with aligned dialogue, people can often infer what is more likely to happen next. Making such predictions requires not only a deep understanding of the rich dynamics…

cs.CV202014 cited

Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models

Jize Cao, Zhe Gan, Yu Cheng +3

Recent Transformer-based large-scale pre-trained models have revolutionized vision-and-language (V+L) research. Models such as ViLBERT, LXMERT and UNITER have significantly lifted…

cs.CV202051 cited

HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Linjie Li, Yen-Chun Chen, Yu Cheng +3

We present HERO, a novel framework for large-scale video+language omni-representation learning. HERO encodes multimodal inputs in a hierarchical structure, where local context of a…

cs.CV2020

BachGAN: High-Resolution Image Synthesis from Salient Object Layout

Yandong Li, Yu Cheng, Zhe Gan +3

We propose a new task towards more practical application for image generation - high-quality image synthesis from salient object layout. This new setting allows users to provide th…

cs.CV2020

VIOLIN: A Large-Scale Dataset for Video-and-Language Inference

Jingzhou Liu, Wenhu Chen, Yu Cheng +4

We introduce a new task, Video-and-Language Inference, for joint multimodal understanding of video and text. Given a video clip with aligned subtitles as premise, paired with a nat…

cs.CV2020

TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval

Jie Lei, Licheng Yu, Tamara L. Berg +1

We introduce TV show Retrieval (TVR), a new multimodal retrieval dataset. TVR requires systems to understand both videos and their associated subtitle (dialogue) texts, making it m…