Reading Wikipedia to Answer Open-Domain Questions
arXiv:1704.00051
Abstract
This paper proposes to tackle open- domain question answering using Wikipedia as the unique knowledge source: the answer to any factoid question is a text span in a Wikipedia article. This task of machine reading at scale combines the challenges of document retrieval (finding the relevant articles) with that of machine comprehension of text (identifying the answer spans from those articles). Our approach combines a search component based on bigram hashing and TF-IDF matching with a multi-layer recurrent neural network model trained to detect answers in Wikipedia paragraphs. Our experiments on multiple existing QA datasets indicate that (1) both modules are highly competitive with respect to existing counterparts and (2) multitask learning using distant supervision on their combination is an effective complete system on this challenging task.
ACL2017, 10 pages
References in corpus (4)
Cited by in corpus (54)
- Recent Trends in Deep Learning Based Natural Language Processing
- BERT Post-Training for Review Reading Comprehension and Aspect-based Sentiment Analysis
- Passage Re-ranking with BERT
- A Survey on Natural Language Processing for Fake News Detection
- Natural Language Inference over Interaction Space
- Personalizing Dialogue Agents: I have a dog, do you have pets too?
- SDNet: Contextualized Attention-based Deep Network for Conversational Question Answering
- A BERT Baseline for the Natural Questions
- Graph Neural Networks for Natural Language Processing: A Survey
- A Unified MRC Framework for Named Entity Recognition
- BAG: Bi-directional Attention Entity Graph Convolutional Network for Multi-hop Reasoning Question Answering
- Technical report on Conversational Question Answering
- Reinforced Mnemonic Reader for Machine Reading Comprehension
- Dice Loss for Data-imbalanced NLP Tasks
- TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
- Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension
- Focused Hierarchical RNNs for Conditional Sequence Processing
- Word2Bits - Quantized Word Vectors
- GraphFlow: Exploiting Conversation Flow with Graph Neural Networks for Conversational Machine Comprehension
- Dynamic Fusion Networks for Machine Reading Comprehension
- Augmenting Transformers with KNN-Based Composite Memory for Dialogue
- Review Conversational Reading Comprehension
- Simple and Effective Curriculum Pointer-Generator Networks for Reading Comprehension over Long Narratives
- Smarnet: Teaching Machines to Read and Comprehend Like Human
- Ranking Paragraphs for Improving Answer Recall in Open-Domain Question Answering
- Yuanfudao at SemEval-2018 Task 11: Three-way Attention and Relational Knowledge for Commonsense Machine Comprehension
- Answering Science Exam Questions Using Query Rewriting with Background Knowledge
- Interpretable Multi-Step Reasoning with Knowledge Extraction on Complex Healthcare Question Answering
- Dimsum @LaySumm 20: BART-based Approach for Scientific Document Summarization
- Explicit Utilization of General Knowledge in Machine Reading Comprehension
- Mining Implicit Relevance Feedback from User Behavior for Web Question Answering
- Multi-task Learning with Sample Re-weighting for Machine Reading Comprehension
- Text Embeddings for Retrieval From a Large Knowledge Base
- SirenLess: reveal the intention behind news
- Baleen: Robust Multi-Hop Reasoning at Scale via Condensed Retrieval
- Multi-Mention Learning for Reading Comprehension with Neural Cascades
- Interactive Teaching for Conversational AI
- Complementary Evidence Identification in Open-Domain Question Answering
- Using Holographically Compressed Embeddings in Question Answering
- A Multi-Resolution Word Embedding for Document Retrieval from Large Unstructured Knowledge Bases
- Conditioning LSTM Decoder and Bi-directional Attention Based Question Answering System
- Answering questions by learning to rank -- Learning to rank by answering questions
- Achieving Human Parity on Visual Question Answering
- Question Answering via Web Extracted Tables and Pipelined Models
- Open-Domain Question Answering with Pre-Constructed Question Spaces
- A Study of the Tasks and Models in Machine Reading Comprehension
- Query-Based Named Entity Recognition
- Towards Language Agnostic Universal Representations
- Efficient Retrieval Optimized Multi-task Learning
- Weighted Global Normalization for Multiple Choice Reading Comprehension over Long Documents
- Decoupled Transformer for Scalable Inference in Open-domain Question Answering
- Learning to Summarize Passages: Mining Passage-Summary Pairs from Wikipedia Revision Histories
- Bew: Towards Answering Business-Entity-Related Web Questions
- ODSQA: Open-domain Spoken Question Answering Dataset