activity
20152022
most citedMake-A-Video: Text-to-Video Generation without Text-Video Data

315 citations · 955 across the 38 of their papers we have counts for

collaborators
Showing 2018Show all

13 papers · 1 filter

cs.CV2018

nocaps: novel object captioning at scale

Harsh Agrawal, Karan Desai, Yufei Wang +7

Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these m…

cs.AI2018

Do Explanations make VQA Models more Predictable to a Human?

Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav +2

A rich line of research attempts to make deep neural networks more transparent by generating human-interpretable 'explanations' of their decision process, especially for interactiv…

cs.LG2018

TarMAC: Targeted Multi-Agent Communication

Abhishek Das, Théophile Gervet, Joshua Romoff +4

We propose a targeted communication architecture for multi-agent reinforcement learning, where agents learn both what messages to send and whom to address them to while performing…

cs.AI2018

Neural Modular Control for Embodied Question Answering

Abhishek Das, Georgia Gkioxari, Stefan Lee +2

We present a modular approach for learning policies for navigation over long planning horizons from language input. Our hierarchical policy operates at multiple timescales, where t…

cs.RO2018

Visual Curiosity: Learning to Ask Questions to Learn Visual Recognition

Jianwei Yang, Jiasen Lu, Stefan Lee +2

In an open-world setting, it is inevitable that an intelligent agent (e.g., a robot) will encounter visual objects, attributes or relationships it does not recognize. In this work,…

cs.CV2018

Visual Coreference Resolution in Visual Dialog using Neural Module Networks

Satwik Kottur, José M. F. Moura, Devi Parikh +2

Visual dialog entails answering a series of questions grounded in an image, using dialog history as context. In addition to the challenges found in visual question answering (VQA),…