activity
20152022
most citedMake-A-Video: Text-to-Video Generation without Text-Video Data

315 citations · 955 across the 38 of their papers we have counts for

collaborators
Showing 2020Show all

21 papers · 1 filter

cs.CV202010 cited

Object-Centric Diagnosis of Visual Reasoning

Jianwei Yang, Jiayuan Mao, Jiajun Wu +4

When answering questions about an image, it not only needs knowing what -- understanding the fine-grained contents (e.g., objects, relationships) in the image, but also telling why…

cs.CV20205 cited

KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQA

Kenneth Marino, Xinlei Chen, Devi Parikh +2

One of the most challenging question types in VQA is when answering the question requires outside knowledge not present in the image. In this work we study open-domain knowledge, t…

cs.CV2020

SOrT-ing VQA Models : Contrastive Gradient Learning for Improved Consistency

Sameer Dharur, Purva Tendulkar, Dhruv Batra +2

Recent research in Visual Question Answering (VQA) has revealed state-of-the-art models to be inconsistent in their understanding of the world -- they answer seemingly difficult qu…

cs.CV202021 cited

Sim-to-Real Transfer for Vision-and-Language Navigation

Peter Anderson, Ayush Shrivastava, Joanne Truong +4

We study the challenging problem of releasing a robot in a previously unseen environment, and having it follow unconstrained natural language navigation instructions. Recent work o…

cs.CV202021 cited

Creative Sketch Generation

Songwei Ge, Vedanuj Goswami, C. Lawrence Zitnick +1

Sketching or doodling is a popular creative activity that people engage in. However, most existing work in automatic sketch understanding or generation has focused on sketches that…

cs.CV2020

Where Are You? Localization from Embodied Dialog

Meera Hahn, Jacob Krantz, Dhruv Batra +4

We present Where Are You? (WAY), a dataset of ~6k dialogs in which two humans -- an Observer and a Locator -- complete a cooperative localization task. The Observer is spawned at r…