collaborators

5 papers

cs.AI2026

Scaling Trends for Lie Detector Oversight in Preference Learning

Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day +2

Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detec…

cs.AI2026

SAE-StatSteer: Statistical Consensus Feature Selection for Optimization-Free Activation Steering of Large Language Models

Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy +3

Activation steering adds a residual-stream direction at inference time, providing lightweight behavioral control without fine-tuning. Sparse autoencoders (SAEs) can make such inter…

cs.CY2025

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models

Ann-Kathrin Dombrowski, Dillon Bowen, Adam Gleave +1

Open-weight large language models (LLMs) unlock huge benefits in innovation, personalization, privacy, and democratization. However, their core advantage - modifiability - opens th…

cs.CY2025

AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations

Dillon Bowen, Ann-Kathrin Dombrowski, Adam Gleave +1

The rapid advancement of AI systems has raised widespread concerns about potential harms of frontier AI systems and the need for responsible evaluation and oversight. In this posit…

cs.LG2025

Representation Engineering: A Top-Down Approach to AI Transparency

Andy Zou, Long Phan, Sarah Chen +18

In this paper, we identify and characterize the emerging area of representation engineering (RepE), an approach to enhancing the transparency of AI systems that draws on insights f…