works on

From the 1 of 8 linked papers with an AI index.

collaborators

8 papers

cs.LG2026

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

Anton de la Fuente, Arthur Conmy

The paper investigates whether lessons learned from supervised fine-tuning (SFT) in alignment training, model organisms, and toy models can be transferred across these domains, dem…

cs.LG2026

How Transparent is DiffusionGemma?

Joshua Engels, Callum McDougall, Bilal Chughtai +13

LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors. However, Diffus…

cs.AI2026

Subliminal Learning Is Steering Vector Distillation

Camila Blank, Agam Bhatia, Senthooran Rajamanoharan +2

Subliminal learning refers to a student language model acquiring a teacher's traits (e.g. a system-prompted preference for owls) when fine-tuned on the teacher's outputs, despite t…

cs.CL2026

How do LLMs Compute Verbal Confidence

Dharshan Kumaran, Arthur Conmy, Federico Barbero +3

Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from black-box models. However, how LLMs in…

cs.LG2026

Simple LLM Baselines are Competitive for Model Diffing

Elias Kempf, Simon Schrodi, Bartosz Cywiński +3

Standard LLM evaluations only test capabilities or dispositions that evaluators designed them for, missing unexpected differences such as behavioral shifts between model revisions…

cs.LG2026

Building Production-Ready Probes For Gemini

János Kramár, Joshua Engels, Zheng Wang +4

Frontier language model capabilities are improving rapidly. We thus need stronger mitigations against bad actors misusing increasingly powerful systems. Prior work has shown that a…