8 citations · 9 across the 5 of their papers we have counts for
6 papers
Perturbation CheckLists for Evaluating NLG Evaluation Metrics
Ananya B. Sai, Tanay Dixit, Dev Yashpal Sheth +2
Natural Language Generation (NLG) evaluation is a multifaceted task requiring assessment of multiple desirable criteria, e.g., fluency, coherency, coverage, relevance, adequacy, ov…
Improving Dialog Evaluation with a Multi-reference Adversarial Dataset and Large Scale Pretraining
Ananya B. Sai, Akash Kumar Mohankumar, Siddhartha Arora +1
There is an increasing focus on model-based dialog evaluation metrics such as ADEM, RUBER, and the more recent BERT-based metrics. These models aim to assign a high score to all re…
A Survey of Evaluation Metrics Used for NLG Systems
Ananya B. Sai, Akash Kumar Mohankumar, Mitesh M. Khapra
The success of Deep Learning has created a surge in interest in a wide a range of Natural Language Generation (NLG) tasks. Deep Learning has not only pushed the state of the art in…
Frustratingly Poor Performance of Reading Comprehension Models on Non-adversarial Examples
Soham Parikh, Ananya B. Sai, Preksha Nema +1
When humans learn to perform a difficult task (say, reading comprehension (RC) over longer passages), it is typically the case that their performance improves significantly on an e…
ElimiNet: A Model for Eliminating Options for Reading Comprehension with Multiple Choice Questions
Soham Parikh, Ananya B. Sai, Preksha Nema +1
The task of Reading Comprehension with Multiple Choice Questions, requires a human (or machine) to read a given passage, question pair and select one of the n given options. The cu…
Re-evaluating ADEM: A Deeper Look at Scoring Dialogue Responses
Ananya B. Sai, Mithun Das Gupta, Mitesh M. Khapra +1
Automatically evaluating the quality of dialogue responses for unstructured domains is a challenging problem. ADEM(Lowe et al. 2017) formulated the automatic evaluation of dialogue…