7 papers
MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models
Victor Ojewale, Ro Encarnación, Suresh Venkatasubramanian +1
Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms…
Bridging Predictions and Interventions: An Integrated Framework for Automated Decision-Systems
Inioluwa Deborah Raji, Lydia T. Liu, Angela Zhou +27
Automated decision systems (ADS) leverage predictions about individual future outcomes to inform consequential decision-making in organizational settings. Across various settings -…
Designing for Doubt: The Case for Informed Abstention in Autonomous Agents
Victor Ojewale, Suresh Venkatasubramanian
As large language models gain tool access and are deployed as autonomous agents capable of editing records, executing transactions, and modifying infrastructure, we still evaluate…
How to Stop Playing Whack-a-Mole: Mapping the Ecosystem of Technologies Facilitating AI-Generated Non-Consensual Intimate Images
Michelle L. Ding, Harini Suresh, Suresh Venkatasubramanian
The last decade has witnessed a rapid advancement of generative AI technology that significantly scaled the accessibility of AI-generated non-consensual intimate images (AIG-NCII),…
Multi-lingual Functional Evaluation for Large Language Models
Victor Ojewale, Inioluwa Deborah Raji, Suresh Venkatasubramanian
Multi-lingual competence in large language models is often evaluated via static data benchmarks such as Belebele, M-MMLU and M-GSM. However, these evaluations often fail to provide…
Audit Trails for Accountability in Large Language Models
Victor Ojewale, Harini Suresh, Suresh Venkatasubramanian
Large language models (LLMs) are increasingly embedded in consequential decisions across healthcare, finance, employment, and public services. Yet accountability remains fragile be…