12 citations · 24 across the 8 of their papers we have counts for
15 papers · 1 filter
Answer Consolidation: Formulation and Benchmarking
Wenxuan Zhou, Qiang Ning, Heba Elfardy +2
Current question answering (QA) systems primarily consider the single-answer scenario, where each question is assumed to be paired with one correct answer. However, in many real-wo…
SpartQA: : A Textual Question Answering Benchmark for Spatial Reasoning
Roshanak Mirzaee, Hossein Rajaby Faghihi, Qiang Ning +1
This paper proposes a question-answering (QA) benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior work and…
ESTER: A Machine Reading Comprehension Dataset for Event Semantic Relation Reasoning
Rujun Han, I-Hung Hsu, Jiao Sun +4
Understanding how events are semantically related to each other is the essence of reading comprehension. Recent event-centric reading comprehension datasets focus mostly on event a…
Temporal Common Sense Acquisition with Minimal Supervision
Ben Zhou, Qiang Ning, Daniel Khashabi +1
Temporal common sense (e.g., duration and frequency of events) is crucial for understanding natural language. However, its acquisition is challenging, partly because such informati…
TORQUE: A Reading Comprehension Dataset of Temporal Ordering Questions
Qiang Ning, Hao Wu, Rujun Han +3
A critical part of reading is being able to understand the temporal relationships between events described in a passage of text, even when those relationships are not explicitly st…
Evaluating Models' Local Decision Boundaries via Contrast Sets
Matt Gardner, Yoav Artzi, Victoria Basmova +23
Standard test sets for supervised learning evaluate in-distribution generalization. Unfortunately, when a dataset has systematic gaps (e.g., annotation artifacts), these evaluation…