5 papers
MoreHopQA: More Than Multi-hop Reasoning
Julian Schnitzler, Xanh Ho, Jiahao Huang +3
Most existing multi-hop datasets are extractive answer datasets, where the answers to the questions can be extracted directly from the provided context. This often leads models to…
What Makes Language Models Good-enough?
Daiki Asami, Saku Sugawara
Psycholinguistic research suggests that humans may build a representation of linguistic input that is 'good-enough' for the task at hand. This study examines what architectural fea…
Probing Physical Reasoning with Counter-Commonsense Context
Kazushi Kondo, Saku Sugawara, Akiko Aizawa
In this study, we create a CConS (Counter-commonsense Contextual Size comparison) dataset to investigate how physical commonsense affects the contextualized size comparison task; t…
On Degrees of Freedom in Defining and Testing Natural Language Understanding
Saku Sugawara, Shun Tsugita
Natural language understanding (NLU) studies often exaggerate or underestimate the capabilities of systems, thereby limiting the reproducibility of their findings. These erroneous…
Analyzing the Effectiveness of the Underlying Reasoning Tasks in Multi-hop Question Answering
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara +1
To explain the predicted answers and evaluate the reasoning abilities of models, several studies have utilized underlying reasoning (UR) tasks in multi-hop question answering (QA)…