8 papers
When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming
Zelalem Abahana, David Evans, Satish Mahadevan Srinivasan +1
RLHF evaluation should track how failures emerge, where they localize, and which warning signals appear before external quality degrades. We study this problem with a compact RLHF…
White-Box Sensitivity Auditing with Steering Vectors
Hannah Cyberey, Yangfeng Ji, David Evans
Algorithmic audits are essential tools for examining systems for properties required by regulators or desired by operators. Current audits of large language models (LLMs) primarily…
Inferring Events from Time Series using Language Models
Mingtian Tan, Mike A. Merrill, Zack Gottesman +3
A common goal in analyzing time series data is to understand how events cause observed variations. We study whether Large Language Models (LLMs) can infer natural language events a…
NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning
Michael Jerge, David Evans
This paper presents NoisyCoconut, a novel inference-time method that enhances large language model (LLM) reliability by manipulating internal representations. Unlike fine-tuning me…
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
Hannah Cyberey, Yangfeng Ji, David Evans
Allocational harms occur when resources or opportunities are unfairly withheld from specific groups. Many proposed bias measures ignore the discrepancy between predictions, which a…
Unsupervised Concept Vector Extraction for Bias Control in LLMs
Hannah Cyberey, Yangfeng Ji, David Evans
Large language models (LLMs) are known to perpetuate stereotypes and exhibit biases. Various strategies have been proposed to mitigate these biases, but most work studies biases as…