A Comparative Study of Word Embeddings for Reading Comprehension
arXiv:1703.00993
Abstract
The focus of past machine learning research for Reading Comprehension tasks has been primarily on the design of novel deep learning architectures. Here we show that seemingly minor choices made on (1) the use of pre-trained word embeddings, and (2) the representation of out-of-vocabulary tokens at test time, can turn out to have a larger impact than architectural choices on the final performance. We systematically explore several options for these choices, and provide recommendations to researchers working in this area.
References in corpus (12)
- Adam: A Method for Stochastic Optimization
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- SQuAD: 100,000+ Questions for Machine Comprehension of Text
- The Goldilocks Principle: Reading Children's Books with Explicit Memory Representations
- Dynamic Coattention Networks For Question Answering
- Machine Comprehension Using Match-LSTM and Answer Pointer
- ReasoNet: Learning to Stop Reading in Machine Comprehension
- A Thorough Examination of the CNN/Daily Mail Reading Comprehension Task
- Multi-Perspective Context Matching for Machine Comprehension
- NewsQA: A Machine Comprehension Dataset
- Embracing data abundance: BookTest Dataset for Reading Comprehension
- End-to-End Answer Chunk Extraction and Ranking for Reading Comprehension
Cited by in corpus (8)
- Learning to Compute Word Embeddings On the Fly
- Neural Machine Reading Comprehension: Methods and Trends
- An Exploration of Word Embedding Initialization in Deep-Learning Tasks
- Exploring Benefits of Transfer Learning in Neural Machine Translation
- Word2Vec: Optimal Hyper-Parameters and Their Impact on NLP Downstream Tasks
- An Investigation of the Interactions Between Pre-Trained Word Embeddings, Character Models and POS Tags in Dependency Parsing
- Enhancing Semantic Word Representations by Embedding Deeper Word Relationships
- Neural Supervised Domain Adaptation by Augmenting Pre-trained Models with Random Units