182 citations · 200 across the 2 of their papers we have counts for
6 papers
Object-Centric Image Generation from Layouts
Tristan Sylvain, Pengchuan Zhang, Yoshua Bengio +2
Despite recent impressive results on single-object and single-domain image generation, the generation of complex scenes with multiple objects remains challenging. In this paper, we…
From FiLM to Video: Multi-turn Question Answering with Multi-modal Context
Dat Tien Nguyen, Shikhar Sharma, Hannes Schulz +1
Understanding audio-visual content and the ability to have an informative conversation about it have both been challenging areas for intelligent systems. The Audio Visual Scene-awa…
Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic Instruction
Alaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz +5
Conditional text-to-image generation is an active area of research, with many possible applications. Existing research has primarily focused on generating a single image from avail…
ChatPainter: Improving Text to Image Generation using Dialogue
Shikhar Sharma, Dendi Suhubdy, Vincent Michalski +2
Synthesizing realistic images from text descriptions on a dataset like Microsoft Common Objects in Context (MS COCO), where each image can contain several objects, is a challenging…
Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation
Shikhar Sharma, Layla El Asri, Hannes Schulz +1
Automated metrics such as BLEU are widely used in the machine translation literature. They have also been used recently in the dialogue community for evaluating dialogue response g…
A Frame Tracking Model for Memory-Enhanced Dialogue Systems
Hannes Schulz, Jeremie Zumer, Layla El Asri +1
Recently, resources and tasks were proposed to go beyond state tracking in dialogue systems. An example is the frame tracking task, which requires recording multiple frames, one fo…