182 citations · 200 across the 2 of their papers we have counts for
3 papers · 1 filter
From FiLM to Video: Multi-turn Question Answering with Multi-modal Context
Dat Tien Nguyen, Shikhar Sharma, Hannes Schulz +1
Understanding audio-visual content and the ability to have an informative conversation about it have both been challenging areas for intelligent systems. The Audio Visual Scene-awa…
Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic Instruction
Alaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz +5
Conditional text-to-image generation is an active area of research, with many possible applications. Existing research has primarily focused on generating a single image from avail…
ChatPainter: Improving Text to Image Generation using Dialogue
Shikhar Sharma, Dendi Suhubdy, Vincent Michalski +2
Synthesizing realistic images from text descriptions on a dataset like Microsoft Common Objects in Context (MS COCO), where each image can contain several objects, is a challenging…