5 papers
White-Box Sensitivity Auditing with Steering Vectors
Hannah Cyberey, Yangfeng Ji, David Evans
Algorithmic audits are essential tools for examining systems for properties required by regulators or desired by operators. Current audits of large language models (LLMs) primarily…
Aligning Language Model Benchmarks with Pairwise Preferences
Marco Gutierrez, Xinyi Leng, Hannah Cyberey +3
Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find that benchmarks often fail to predict real…
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
Hannah Cyberey, Yangfeng Ji, David Evans
Allocational harms occur when resources or opportunities are unfairly withheld from specific groups. Many proposed bias measures ignore the discrepancy between predictions, which a…
Unsupervised Concept Vector Extraction for Bias Control in LLMs
Hannah Cyberey, Yangfeng Ji, David Evans
Large language models (LLMs) are known to perpetuate stereotypes and exhibit biases. Various strategies have been proposed to mitigate these biases, but most work studies biases as…
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
Hannah Cyberey, David Evans
Large language models (LLMs) have transformed the way we access information. These models are often tuned to refuse to comply with requests that are considered harmful and to produ…