85 citations · 241 across the 6 of their papers we have counts for
3 papers · 1 filter
Towards VQA Models That Can Read
Amanpreet Singh, Vivek Natarajan, Meet Shah +5
Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models ca…
Visual Storytelling
Ting-Hao, Huang, Francis Ferraro +13
We introduce the first dataset for sequential vision-to-language, and explore how this data may be used for the task of visual storytelling. The first release of this dataset, SIND…
A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He +5
Representation and learning of commonsense knowledge is one of the foundational problems in the quest to enable deep language understanding. This issue is particularly challenging…