A Decomposable Attention Model for Natural Language Inference
arXiv:1606.01933
Abstract
We propose a simple neural architecture for natural language inference. Our approach uses attention to decompose the problem into subproblems that can be solved separately, thus making it trivially parallelizable. On the Stanford Natural Language Inference (SNLI) dataset, we obtain state-of-the-art results with almost an order of magnitude fewer parameters than previous work and without relying on any word-order information. Adding intra-sentence attention that takes a minimum amount of order into account yields further improvements.
7 pages, 1 figure, Proceeedings of EMNLP 2016
References in corpus (5)
- Convolutional Neural Network Architectures for Matching Natural Language Sentences
- A large annotated corpus for learning natural language inference
- Long Short-Term Memory-Networks for Machine Reading
- Order-Embeddings of Images and Language
- Natural Language Inference by Tree-Based Convolution and Heuristic Matching
Cited by in corpus (43)
- Deep Learning Based Text Classification: A Comprehensive Review
- Weighted Transformer Network for Machine Translation
- Bilateral Multi-Perspective Matching for Natural Language Sentences
- An Empirical Study of Spatial Attention Mechanisms in Deep Networks
- A Survey of Deep Learning Techniques for Neural Machine Translation
- Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning
- A BERT Baseline for the Natural Questions
- Frame-Semantic Parsing with Softmax-Margin Segmental RNNs and a Syntactic Scaffold
- Deep Learning Based Chatbot Models
- Enhance word representation for out-of-vocabulary on Ubuntu dialogue corpus
- Style Mixer: Semantic-aware Multi-Style Transfer Network
- KG^2: Learning to Reason Science Exam Questions with Contextual Knowledge Graph Embeddings
- On the Importance of Delexicalization for Fact Verification
- Read + Verify: Machine Reading Comprehension with Unanswerable Questions
- Lightweight and Efficient Neural Natural Language Processing with Quaternion Networks
- Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA
- Improving Natural Language Inference Using External Knowledge in the Science Questions Domain
- Knowledge Enhanced Attention for Robust Natural Language Inference
- A survey of Community Question Answering
- Dynamic Multi-Level Multi-Task Learning for Sentence Simplification
- A Neural Architecture Mimicking Humans End-to-End for Natural Language Inference
- Sanity Check: A Strong Alignment and Information Retrieval Baseline for Question Answering
- Long Short-Term Attention
- Detecting and Explaining Causes From Text For a Time Series Event
- Adversarial Training for Community Question Answer Selection Based on Multi-scale Matching
- Path-Based Contextualization of Knowledge Graphs for Textual Entailment
- Interpretation of Natural Language Rules in Conversational Machine Reading
- Multi-Mention Learning for Reading Comprehension with Neural Cascades
- Endowing Deep 3D Models with Rotation Invariance Based on Principal Component Analysis
- Attention Boosted Sequential Inference Model
- Interpretable Self-Attention Temporal Reasoning for Driving Behavior Understanding
- Explaining Neural Network Predictions on Sentence Pairs via Learning Word-Group Masks
- ASBERT: Siamese and Triplet network embedding for open question answering
- Transformer for Emotion Recognition
- AWE: Asymmetric Word Embedding for Textual Entailment
- Abductive Reasoning as Self-Supervision for Common Sense Question Answering
- Recognising Agreement and Disagreement between Stances with Reason Comparing Networks
- Dropping Networks for Transfer Learning
- Less Memory, Faster Speed: Refining Self-Attention Module for Image Reconstruction
- POG: Personalized Outfit Generation for Fashion Recommendation at Alibaba iFashion
- Exploiting Inter-pixel Correlations in Unsupervised Domain Adaptation for Semantic Segmentation
- Multi-Perspective Inferrer: Reasoning Sentences Relationship from Holistic Perspective
- Two-Stream Appearance Transfer Network for Person Image Generation