3 papers
cs.CL2026
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge
Karen Zhou, Chenhao Tan
Checklists have emerged as a popular approach for interpretable and fine-grained evaluation, particularly with LLM-as-a-Judge. Beyond evaluation, these structured criteria can serv…
cs.CL2025
From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes
Karen Zhou, John Giorgi, Pranav Mani +3
AI-generated clinical notes are increasingly used in healthcare, but evaluating their quality remains a challenge due to high subjectivity and limited scalability of expert review.…
cs.CL2025
Quantifying the Uniqueness and Divisiveness of Presidential Discourse
Karen Zhou, Alexander A. Meitus, Milo Chase +4
Do American presidents speak discernibly different from each other? If so, in what ways? And are these differences confined to any single medium of communication? To investigate th…