5 papers
Removing Sandbagging in LLMs by Training with Weak Supervision
Emil Ryd, Henning Bartsch, Julian Stastny +2
As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality. A model more cap…
Eliciting Secret Knowledge from Language Models
Bartosz CywiÅski, Emil Ryd, Rowan Wang +4
We study secret elicitation: discovering knowledge that an AI possesses but does not explicitly verbalize. As a testbed, we train three families of large language models (LLMs) to…
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
Nevan Wichers, Aram Ebtekar, Ariana Azarbal +8
Large language models are sometimes trained with imperfect oversight signals, leading to undesired behaviors such as reward hacking and sycophancy. Improving oversight quality can…
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
Bartosz CywiÅski, Emil Ryd, Senthooran Rajamanoharan +1
As language models become more powerful and sophisticated, it is crucial that they remain trustworthy and reliable. There is concerning preliminary evidence that models may attempt…
Fine Flood Forecasts: Incorporating local data into global models through fine-tuning
Emil Ryd, Grey Nearing
Floods are the most common form of natural disaster and accurate flood forecasting is essential for early warning systems. Previous work has shown that machine learning (ML) models…