Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
arXiv:1706.01554
Abstract
We present a novel training framework for neural sequence models, particularly for grounded dialog generation. The standard training paradigm for these models is maximum likelihood estimation (MLE), or minimizing the cross-entropy of the human responses. Across a variety of domains, a recurring problem with MLE trained generative neural dialog models (G) is that they tend to produce 'safe' and generic responses ("I don't know", "I can't tell"). In contrast, discriminative dialog models (D) that are trained to rank a list of candidate human responses outperform their generative counterparts; in terms of automatic metrics, diversity, and informativeness of the responses. However, D is not useful in practice since it cannot be deployed to have real conversations with users. Our work aims to achieve the best of both worlds -- the practical usefulness of G and the strong performance of D -- via knowledge transfer from D to G. Our primary contribution is an end-to-end trainable generative visual dialog model, where G receives gradients from D as a perceptual (not adversarial) loss of the sequence sampled from G. We leverage the recently proposed Gumbel-Softmax (GS) approximation to the discrete distribution -- specifically, an RNN augmented with a sequence of GS samplers, coupled with the straight-through gradient estimator to enable end-to-end differentiability. We also introduce a stronger encoder for visual dialog, and employ a self-attention mechanism for answer encoding along with a metric learning loss to aid D in better capturing semantic similarities in answer responses. Overall, our proposed model outperforms state-of-the-art on the VisDial dataset by a significant margin (2.67% on recall@10). The source code can be downloaded from https://github.com/jiasenlu/visDial.pytorch.
11 pages, 3 figures
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- Energy-based Generative Adversarial Network
- GANS for Sequences of Discrete Elements with the Gumbel-softmax Distribution
- Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning
- A Neural Network Approach to Context-Sensitive Generation of Conversational Responses
- End-to-end optimization of goal-driven and visually grounded dialogue systems
- GuessWhat?! Visual object discovery through multi-modal dialogue
Cited by in corpus (6)
- Asking the Difficult Questions: Goal-Oriented Visual Question Generation via Intermediate Rewards
- Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
- Learning to Respond with Your Favorite Stickers: A Framework of Unifying Multi-Modality and User Preference in Multi-Turn Dialog
- DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue
- Dialog without Dialog Data: Learning Visual Dialog Agents from VQA Data
- ORD: Object Relationship Discovery for Visual Dialogue Generation