4 papers
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
Charles Chiang, Simret Gebreegziabher, Annalisa Szymanski +6
LLM-as-a-judge approaches have emerged as a scalable solution for evaluating model behaviors, yet they rely on evaluation criteria often created by a single individual, embedding t…
Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic Study
Annalisa Szymanski, Oghenemaro Anuyah, Toby Jia-Jun Li +1
Large Language Models (LLMs) are increasingly developed for use in complex professional domains, yet little is known about how teams design and evaluate these systems in practice.…
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
Annalisa Szymanski, Simret Araya Gebreegziabher, Oghenemaro Anuyah +2
Large Language Models (LLMs) are increasingly utilized for domain-specific tasks, yet evaluating their outputs remains challenging. A common strategy is to apply evaluation criteri…
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
Annalisa Szymanski, Noah Ziems, Heather A. Eicher-Miller +3
The potential of using Large Language Models (LLMs) themselves to evaluate LLM outputs offers a promising method for assessing model performance across various contexts. Previous r…