7 citations · 15 across the 3 of their papers we have counts for
3 papers
DQI: A Guide to Benchmark Evaluation
Swaroop Mishra, Anjana Arunkumar, Bhavdeep Sachdeva +2
A `state of the art' model A surpasses humans in a benchmark B, but fails on similar benchmarks C, D, and E. What does B have that the other benchmarks do not? Recent research prov…
Our Evaluation Metric Needs an Update to Encourage Generalization
Swaroop Mishra, Anjana Arunkumar, Chris Bryan +1
Models that surpass human performance on several popular benchmarks display significant degradation in performance on exposure to Out of Distribution (OOD) data. Recent research ha…
DQI: Measuring Data Quality in NLP
Swaroop Mishra, Anjana Arunkumar, Bhavdeep Sachdeva +2
Neural language models have achieved human level performance across several NLP datasets. However, recent studies have shown that these models are not truly learning the desired ta…