activity
20212026
most citedTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

391 citations · 551 across the 51 of their papers we have counts for

collaborators
Showing cs.LGShow all

34 papers · 1 filter

cs.LG2026

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

Aditya Singh, Gerson Kroiz, Senthooran Rajamanoharan +1

A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior. But behavior alone does not establi…

cs.LG2026

How Transparent is DiffusionGemma?

Joshua Engels, Callum McDougall, Bilal Chughtai +11

LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors. However, Diffus…

cs.LG2026

Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation

Helena Casademunt, Bartosz Cywiński, Khoi Tran +3

Large language models sometimes produce false or misleading responses. Two approaches to this problem are honesty elicitation -- modifying prompts or weights so that the model answ…

cs.LG2026

Simple LLM Baselines are Competitive for Model Diffing

Elias Kempf, Simon Schrodi, Bartosz Cywiński +3

Standard LLM evaluations only test capabilities or dispositions that evaluators designed them for, missing unexpected differences such as behavioral shifts between model revisions…

cs.LG2026

Building Production-Ready Probes For Gemini

János Kramár, Joshua Engels, Zheng Wang +4

Frontier language model capabilities are improving rapidly. We thus need stronger mitigations against bad actors misusing increasingly powerful systems. Prior work has shown that a…

cs.LG2025

Difficulties with Evaluating a Deception Detector for AIs

Lewis Smith, Bilal Chughtai, Neel Nanda

Building reliable deception detectors for AI systems -- methods that could predict when an AI system is being strategically deceptive without necessarily requiring behavioural evid…