activity
20152022
most citedMake-A-Video: Text-to-Video Generation without Text-Video Data

315 citations · 955 across the 38 of their papers we have counts for

collaborators
Showing 2016Show all

5 papers · 1 filter

stat.ML2016

Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?

Abhishek Das, Harsh Agrawal, C. Lawrence Zitnick +2

We conduct large-scale studies on `human attention' in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and tes…

cs.CV2016

Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?

Abhishek Das, Harsh Agrawal, C. Lawrence Zitnick +2

We conduct large-scale studies on `human attention' in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and tes…

cs.CL2016

Visual Storytelling

Ting-Hao, Huang, Francis Ferraro +13

We introduce the first dataset for sequential vision-to-language, and explore how this data may be used for the task of visual storytelling. The first release of this dataset, SIND…

cs.CL2016

A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories

Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He +5

Representation and learning of commonsense knowledge is one of the foundational problems in the quest to enable deep language understanding. This issue is particularly challenging…

cs.CV2016

Joint Unsupervised Learning of Deep Representations and Image Clusters

Jianwei Yang, Devi Parikh, Dhruv Batra

In this paper, we propose a recurrent framework for Joint Unsupervised LEarning (JULE) of deep representations and image clusters. In our framework, successive operations in a clus…