Reasoning About Pragmatics with Neural Listeners and Speakers
arXiv:1604.00562
Abstract
We present a model for pragmatically describing scenes, in which contrastive behavior results from a combination of inference-driven pragmatics and learned semantics. Like previous learned approaches to language generation, our model uses a simple feature-driven architecture (here a pair of neural "listener" and "speaker" models) to ground language in the world. Like inference-driven approaches to pragmatics, our model actively reasons about listener behavior when selecting utterances. For training, our approach requires only ordinary captions, annotated _without_ demonstration of the pragmatic behavior the model ultimately exhibits. In human evaluations on a referring expression game, our approach succeeds 81% of the time, compared to a 69% success rate using existing techniques.
Cited by in corpus (21)
- Discriminability objective for training descriptive captions
- It Takes Two to Tango: Towards Theory of AI's Mind
- Speaking the Same Language: Matching Machine to Human Captions by Adversarial Training
- Emergent Communication in a Multi-Modal, Multi-Step Referential Game
- Context-aware Captions from Context-agnostic Supervision
- A Paradigm for Situated and Goal-Driven Language Learning
- Emergence of Pragmatics from Referential Game between Theory of Mind Agents
- Unified Pragmatic Models for Generating and Following Instructions
- Colors in Context: A Pragmatic Neural Model for Grounded Language Understanding
- A Joint Speaker-Listener-Reinforcer Model for Referring Expressions
- Interactive Learning from Activity Description
- A perspective on multi-agent communication for information fusion
- Context-Aware Group Captioning via Self-Attention and Contrastive Features
- Reasoning about Fine-grained Attribute Phrases using Reference Games
- Learning to refer informatively by amortizing pragmatic reasoning
- Evaluating Text-to-Image Matching using Binary Image Selection (BISON)
- Tell-the-difference: Fine-grained Visual Descriptor via a Discriminating Referee
- Grounding Visual Explanations
- Informative Object Annotations: Tell Me Something I Don't Know
- Planning, Inference and Pragmatics in Sequential Language Games
- Mapping Natural Language Commands to Web Elements