From the 1 of 8 linked papers with an AI index.
8 papers
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
Anton de la Fuente, Arthur Conmy
The paper investigates whether lessons learned from supervised fine-tuning (SFT) in alignment training, model organisms, and toy models can be transferred across these domains, dem…
How Transparent is DiffusionGemma?
Joshua Engels, Callum McDougall, Bilal Chughtai +13
LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors. However, Diffus…
Subliminal Learning Is Steering Vector Distillation
Camila Blank, Agam Bhatia, Senthooran Rajamanoharan +2
Subliminal learning refers to a student language model acquiring a teacher's traits (e.g. a system-prompted preference for owls) when fine-tuned on the teacher's outputs, despite t…
How do LLMs Compute Verbal Confidence
Dharshan Kumaran, Arthur Conmy, Federico Barbero +3
Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from black-box models. However, how LLMs in…
Simple LLM Baselines are Competitive for Model Diffing
Elias Kempf, Simon Schrodi, Bartosz CywiÅski +3
Standard LLM evaluations only test capabilities or dispositions that evaluators designed them for, missing unexpected differences such as behavioral shifts between model revisions…
Building Production-Ready Probes For Gemini
János Kramár, Joshua Engels, Zheng Wang +4
Frontier language model capabilities are improving rapidly. We thus need stronger mitigations against bad actors misusing increasingly powerful systems. Prior work has shown that a…