2 papers
cs.CL2025
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
Vipul Gupta, Candace Ross, David Pantoja +3
One of the most challenging problems facing NLP today is evaluation. Some of the most pressing issues pertain to benchmark saturation, data contamination, and diversity in the qual…
cs.CL2024
Changing Answer Order Can Decrease MMLU Accuracy
Vipul Gupta, David Pantoja, Candace Ross +2
As large language models (LLMs) have grown in prevalence, particular benchmarks have become essential for the evaluation of these models and for understanding model capabilities. M…