384 citations · 1.3k across the 44 of their papers we have counts for
7 papers · 1 filter
Towards VQA Models That Can Read
Amanpreet Singh, Vivek Natarajan, Meet Shah +5
Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models ca…
Dialog System Technology Challenge 7
Koichiro Yoshino, Chiori Hori, Julien Perez +14
This paper introduces the Seventh Dialog System Technology Challenges (DSTC), which use shared datasets to explore the problem of building dialog systems. Recently, end-to-end dial…
End-to-End Audio Visual Scene-Aware Dialog using Multimodal Attention-Based Video Features
Chiori Hori, Huda Alamri, Jue Wang +10
Dialog systems need to understand dynamic visual scenes in order to have conversations with users about the objects and events around them. Scene-aware dialog systems for real-worl…
Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7
Huda Alamri, Vincent Cartillier, Raphael Gontijo Lopes +8
Scene-aware dialog systems will be able to have conversations with users about the objects and events around them. Progress on such systems can be made by integrating state-of-the-…
Natural Language Does Not Emerge 'Naturally' in Multi-Agent Dialog
Satwik Kottur, José M. F. Moura, Stefan Lee +1
A number of recent works have proposed techniques for end-to-end learning of communication protocols among cooperative multi-agent populations, and have simultaneously found the em…
Visual Storytelling
Ting-Hao, Huang, Francis Ferraro +13
We introduce the first dataset for sequential vision-to-language, and explore how this data may be used for the task of visual storytelling. The first release of this dataset, SIND…