activity
20232026
collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2025

Doubly-Robust LLM-as-a-Judge: Externally Valid Estimation with Imperfect Personas

Luke Guerdan, Justin Whitehouse, Kimberly Truong +2

As Generative AI (GenAI) systems see growing adoption, a key concern involves the external validity of evaluations, or the extent to which they generalize from lab-based to real-wo…

cs.LG2025

Bridging Prediction and Intervention Problems in Social Systems

Lydia T. Liu, Inioluwa Deborah Raji, Angela Zhou +32

Many automated decision systems (ADS) are designed to solve prediction problems -- where the goal is to learn patterns from a sample of the population and apply them to individuals…

cs.LG2025

Validating LLM-as-a-Judge Systems under Rating Indeterminacy

Luke Guerdan, Solon Barocas, Kenneth Holstein +3

The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and st…

cs.LG2024

A Framework for Evaluating LLMs Under Task Indeterminacy

Luke Guerdan, Hanna Wallach, Solon Barocas +1

Large language model (LLM) evaluations often assume there is a single correct response -- a gold label -- for each item in the evaluation corpus. However, some tasks can be ambiguo…

cs.LG2024

Predictive Performance Comparison of Decision Policies Under Confounding

Luke Guerdan, Amanda Coston, Kenneth Holstein +1

Predictive models are often introduced to decision-making tasks under the rationale that they improve performance over an existing decision-making policy. However, it is challengin…