papers

Publications (5)

cs.CL2022

NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks

Swaroop Mishra, Arindam Mitra, Neeraj Varshney +4

Given the ubiquitous nature of numbers in text, reasoning with numbers to perform simple calculations is an important skill of AI systems. While many datasets and models have been…

cs.CL2020

Towards Question Format Independent Numerical Reasoning: A Set of Prerequisite Tasks

Swaroop Mishra, Arindam Mitra, Neeraj Varshney +2

Numerical reasoning is often important to accurately understand the world. Recently, several format-specific datasets have been proposed, such as numerical reasoning in the setting…

cs.CL2020

DQI: Measuring Data Quality in NLP

Swaroop Mishra, Anjana Arunkumar, Bhavdeep Sachdeva +2

Neural language models have achieved human level performance across several NLP datasets. However, recent studies have shown that these models are not truly learning the desired ta…

cs.CL2023

Real-Time Visual Feedback to Guide Benchmark Creation: A Human-and-Metric-in-the-Loop Workflow

Anjana Arunkumar, Swaroop Mishra, Bhavdeep Sachdeva +2

Recent research has shown that language models exploit `artifacts' in benchmarks to solve tasks, rather than truly learning them, leading to inflated model performance. In pursuit…

cs.CL2020

DQI: A Guide to Benchmark Evaluation

Swaroop Mishra, Anjana Arunkumar, Bhavdeep Sachdeva +2

A `state of the art' model A surpasses humans in a benchmark B, but fails on similar benchmarks C, D, and E. What does B have that the other benchmarks do not? Recent research prov…