collaborators

8 papers

cs.LG2026

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming

Zelalem Abahana, David Evans, Satish Mahadevan Srinivasan +1

RLHF evaluation should track how failures emerge, where they localize, and which warning signals appear before external quality degrades. We study this problem with a compact RLHF…

cs.CY2026

White-Box Sensitivity Auditing with Steering Vectors

Hannah Cyberey, Yangfeng Ji, David Evans

Algorithmic audits are essential tools for examining systems for properties required by regulators or desired by operators. Current audits of large language models (LLMs) primarily…

cs.AI2026

Inferring Events from Time Series using Language Models

Mingtian Tan, Mike A. Merrill, Zack Gottesman +3

A common goal in analyzing time series data is to understand how events cause observed variations. We study whether Large Language Models (LLMs) can infer natural language events a…

cs.LG2026

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning

Michael Jerge, David Evans

This paper presents NoisyCoconut, a novel inference-time method that enhances large language model (LLM) reliability by manipulating internal representations. Unlike fine-tuning me…

cs.CL2026

Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?

Hannah Cyberey, Yangfeng Ji, David Evans

Allocational harms occur when resources or opportunities are unfairly withheld from specific groups. Many proposed bias measures ignore the discrepancy between predictions, which a…

cs.CL2025

Unsupervised Concept Vector Extraction for Bias Control in LLMs

Hannah Cyberey, Yangfeng Ji, David Evans

Large language models (LLMs) are known to perpetuate stereotypes and exhibit biases. Various strategies have been proposed to mitigate these biases, but most work studies biases as…