Identifying Well-formed Natural Language Questions
arXiv:1808.09419
Abstract
Understanding search queries is a hard problem as it involves dealing with "word salad" text ubiquitously issued by users. However, if a query resembles a well-formed question, a natural language processing pipeline is able to perform more accurate interpretation, thus reducing downstream compounding errors. Hence, identifying whether or not a query is well formed can enhance query understanding. Here, we introduce a new task of identifying a well-formed natural language question. We construct and release a dataset of 25,100 publicly available questions classified into well-formed and non-wellformed categories and report an accuracy of 70.7% on the test set. We also show that our classifier can be used to improve the performance of neural sequence-to-sequence models for generating questions for reading comprehension.
References in corpus (5)
- Microsoft COCO Captions: Data Collection and Evaluation Server
- SQuAD: 100,000+ Questions for Machine Comprehension of Text
- Ask the Right Questions: Active Question Reformulation with Reinforcement Learning
- Bilateral Multi-Perspective Matching for Natural Language Sentences
- Learning to Ask: Neural Question Generation for Reading Comprehension
Cited by in corpus (4)
- Unsupervised Question Answering by Cloze Translation
- Keyword Extraction for Improved Document Retrieval in Conversational Search
- How to Ask Better Questions? A Large-Scale Multi-Domain Dataset for Rewriting Ill-Formed Questions
- Generative Question Refinement with Deep Reinforcement Learning in Retrieval-based QA System