7 citations · 17 across the 6 of their papers we have counts for
5 papers · 1 filter
Hardness of Samples Need to be Quantified for a Reliable Evaluation System: Exploring Potential Opportunities with a New Task
Swaroop Mishra, Anjana Arunkumar, Chris Bryan +1
Evaluation of models on benchmarks is unreliable without knowing the degree of sample hardness; this subsequently overestimates the capability of AI systems and limits their adopti…
A Survey of Parameters Associated with the Quality of Benchmarks in NLP
Swaroop Mishra, Anjana Arunkumar, Chris Bryan +1
Several benchmarks have been built with heavy investment in resources to track our progress in NLP. Thousands of papers published in response to those benchmarks have competed to t…
DQI: A Guide to Benchmark Evaluation
Swaroop Mishra, Anjana Arunkumar, Bhavdeep Sachdeva +2
A `state of the art' model A surpasses humans in a benchmark B, but fails on similar benchmarks C, D, and E. What does B have that the other benchmarks do not? Recent research prov…
Our Evaluation Metric Needs an Update to Encourage Generalization
Swaroop Mishra, Anjana Arunkumar, Chris Bryan +1
Models that surpass human performance on several popular benchmarks display significant degradation in performance on exposure to Out of Distribution (OOD) data. Recent research ha…
DQI: Measuring Data Quality in NLP
Swaroop Mishra, Anjana Arunkumar, Bhavdeep Sachdeva +2
Neural language models have achieved human level performance across several NLP datasets. However, recent studies have shown that these models are not truly learning the desired ta…