7 papers
Online Safety Monitoring for LLMs
Mona Schirmer, Metod Jazbec, Alexander Timans +3
Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed i…
Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting
Andrea Wynn, Metod Jazbec, Charith Peris +4
Large language models (LLMs) can be influenced by harmful or irrelevant context, which can significantly harm model performance on downstream tasks. This motivates principled desig…
Conformal Thinking: Risk Control for Reasoning on a Compute Budget
Xi Wang, Anushri Suresh, Alvin Zhang +6
Reasoning Large Language Models (LLMs) enable test-time scaling, with dataset-level accuracy improving as the token budget increases, motivating adaptive reasoning -- spending toke…
Temporal Test-Time Adaptation with State-Space Models
Mona Schirmer, Dan Zhang, Eric Nalisnick
Distribution shifts between training and test data are inevitable over the lifecycle of a deployed model, leading to performance decay. Adapting a model on test samples can help mi…
Monitoring Risks in Test-Time Adaptation
Mona Schirmer, Metod Jazbec, Christian A. Naesseth +1
Encountering shifted data at test time is a ubiquitous challenge when deploying predictive models. Test-time adaptation (TTA) methods address this issue by continuously adapting a…
Generative Uncertainty in Diffusion Models
Metod Jazbec, Eliot Wong-Toi, Guoxuan Xia +3
Diffusion models have recently driven significant breakthroughs in generative modeling. While state-of-the-art models produce high-quality samples on average, individual samples ca…