collaborators

6 papers

cs.CL2026

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis

May Lynn Reese, Markela Zeneli, Mindy Ng +3

General-purpose Large Language Models (LLMs) are becoming widely adopted by people for mental health support. Yet emerging evidence suggests there are significant risks associated…

cs.AI2025

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models

Cameron Tice, Philipp Alexander Kreer, Nathan Helm-Burger +7

Capability evaluations play a crucial role in assessing and regulating frontier AI systems. The effectiveness of these evaluations faces a significant challenge: strategic underper…

cs.AI2025

Approximating Human Preferences Using a Multi-Judge Learned System

Eitán Sprejer, Fernando Avalos, Augusto Bernardi +3

Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Ove…

cs.CY2025

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants

Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes +1

As humans delegate more tasks and decisions to artificial intelligence (AI), we risk losing control of our individual and collective futures. Relatively simple algorithmic systems…

cs.CV2025

Evaluating Precise Geolocation Inference Capabilities of Vision Language Models

Neel Jay, Hieu Minh Nguyen, Trung Dung Hoang +1

The prevalence of Vision-Language Models (VLMs) raises important questions about privacy in an era where visual information is increasingly available. While foundation VLMs demonst…

cs.CL2025

Tailored Truths: Optimizing LLM Persuasion with Personalization and Fabricated Statistics

Jasper Timm, Chetan Talele, Jacob Haimes

Large Language Models (LLMs) are becoming increasingly persuasive, demonstrating the ability to personalize arguments in conversation with humans by leveraging their personal data.…