3 papers
cs.HC2026
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
Charles Chiang, Simret Gebreegziabher, Annalisa Szymanski +6
LLM-as-a-judge approaches have emerged as a scalable solution for evaluating model behaviors, yet they rely on evaluation criteria often created by a single individual, embedding t…
cs.HC2026
Crepe: A Mobile Screen Data Collector Using Graph Query
Yuwen Lu, Meng Chen, Qi Zhao +6
Collecting mobile datasets remains challenging for academic researchers due to limited data access and technical barriers. Commercial organizations often possess exclusive access t…
cs.HC2026
Designing Staged Evaluation Workflows for LLMs: Integrating Domain Experts, Lay Users, and Model-Generated Evaluation Criteria
Annalisa Szymanski, Simret Araya Gebreegziabher, Oghenemaro Anuyah +2
Large Language Models (LLMs) are increasingly utilized for domain-specific tasks, yet evaluating their outputs remains challenging. A common strategy is to apply evaluation criteri…