10 papers
Aggregated Individual Reporting for Post-Deployment Evaluation
Jessica Dai, Inioluwa Deborah Raji, Benjamin Recht +1
The need for developing model evaluations beyond static benchmarking, especially in the post-deployment phase, is now well-understood. At the same time, concerns about the concentr…
Bridging Predictions and Interventions: An Integrated Framework for Automated Decision-Systems
Inioluwa Deborah Raji, Lydia T. Liu, Angela Zhou +27
Automated decision systems (ADS) leverage predictions about individual future outcomes to inform consequential decision-making in organizational settings. Across various settings -…
Multi-lingual Functional Evaluation for Large Language Models
Victor Ojewale, Inioluwa Deborah Raji, Suresh Venkatasubramanian
Multi-lingual competence in large language models is often evaluated via static data benchmarks such as Belebele, M-MMLU and M-GSM. However, these evaluations often fail to provide…
Evaluating Prediction-based Interventions with Human Decision Makers In Mind
Inioluwa Deborah Raji, Lydia Liu
Automated decision systems (ADS) are broadly deployed to inform and support human decision-making across a wide range of consequential settings. However, various context-specific d…
Bridging Prediction and Intervention Problems in Social Systems
Lydia T. Liu, Inioluwa Deborah Raji, Angela Zhou +32
Many automated decision systems (ADS) are designed to solve prediction problems -- where the goal is to learn patterns from a sample of the population and apply them to individuals…
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…