3 papers
cs.CL2026
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
Jessica Huynh, Alfredo Gomez, Athiya Deviyani +3
Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation. However, there is limited statistical analysis of how modifica…
cs.CV2024
Offline Evaluation of Set-Based Text-to-Image Generation
Negar Arabzadeh, Fernando Diaz, Junfeng He
Text-to-Image (TTI) systems often support people during ideation, the early stages of a creative process when exposure to a broad set of relevant images can help explore the design…
cs.IR2024
Pessimistic Evaluation
Fernando Diaz
Traditional evaluation of information access systems has focused primarily on average utility across a set of information needs (information retrieval) or users (recommender system…