3 papers
cs.LG2026
Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas
Luke Guerdan, Justin Whitehouse, Kimberly Truong +2
As Generative AI (GenAI) systems see growing adoption, a key concern involves the external validity of evaluations, or the extent to which they generalize from lab-based to real-wo…
cs.LG2025
Validating LLM-as-a-Judge Systems under Rating Indeterminacy
Luke Guerdan, Solon Barocas, Kenneth Holstein +3
The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and st…
cs.HC2025
Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks
Luke Guerdan, Devansh Saxena, Stevie Chancellor +2
Data scientists often formulate predictive modeling tasks involving fuzzy, hard-to-define concepts, such as the "authenticity" of student writing or the "healthcare need" of a pati…