2 papers
cs.HC2026
Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic Study
Annalisa Szymanski, Oghenemaro Anuyah, Toby Jia-Jun Li +1
Large Language Models (LLMs) are increasingly developed for use in complex professional domains, yet little is known about how teams design and evaluate these systems in practice.…
cs.HC2024
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
Annalisa Szymanski, Noah Ziems, Heather A. Eicher-Miller +3
The potential of using Large Language Models (LLMs) themselves to evaluate LLM outputs offers a promising method for assessing model performance across various contexts. Previous r…