83 citations · 83 across the 1 of their papers we have counts for
6 papers
Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?
Abhishek Das, Harsh Agrawal, C. Lawrence Zitnick +2
We conduct large-scale studies on `human attention' in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and tes…
Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?
Abhishek Das, Harsh Agrawal, C. Lawrence Zitnick +2
We conduct large-scale studies on `human attention' in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and tes…
Visual Storytelling
Ting-Hao, Huang, Francis Ferraro +13
We introduce the first dataset for sequential vision-to-language, and explore how this data may be used for the task of visual storytelling. The first release of this dataset, SIND…
A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He +5
Representation and learning of commonsense knowledge is one of the foundational problems in the quest to enable deep language understanding. This issue is particularly challenging…
Joint Unsupervised Learning of Deep Representations and Image Clusters
Jianwei Yang, Devi Parikh, Dhruv Batra
In this paper, we propose a recurrent framework for Joint Unsupervised LEarning (JULE) of deep representations and image clusters. In our framework, successive operations in a clus…
WhittleSearch: Interactive Image Search with Relative Attribute Feedback
Adriana Kovashka, Devi Parikh, Kristen Grauman
We propose a novel mode of feedback for image search, where a user describes which properties of exemplar images should be adjusted in order to more closely match his/her mental mo…