Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
Jessica Huynh, Alfredo Gomez, Athiya Deviyani +3
Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation. However, there is limited statistical analysis of how modifica…
cs.CL2025
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy
Athiya Deviyani, Fernando Diaz
Meta-evaluation of automatic evaluation metrics -- assessing evaluation metrics themselves -- is crucial for accurately benchmarking natural language processing systems and has imp…