activity
20242026
collaborators

11 papers

cs.AI2026

Online Safety Monitoring for LLMs

Mona Schirmer, Metod Jazbec, Alexander Timans +3

Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed i…

cs.LG2026

Uncertainty Estimation for Molecular Diffusion Models

Paul Seij, Christian A. Naesseth, Stephan Mandt +1

Diffusion models have seen wide adoption for 3D molecular generation, yet they offer no principled signal of when a generated molecule is likely to be of low quality. We propose a…

cs.LG2026

Re-evaluating Confidence Remasking in Masked Diffusion Language Models

Stipe Frkovic, Metod Jazbec, Dan Zhang +3

Masked diffusion language models (dLLMs) have recently emerged as a competitive alternative to autoregressive language models, with the promise of faster inference via parallel tok…

cs.LG2026

Learning Unmasking Policies for Diffusion Language Models

Metod Jazbec, Theo X. Olausson, Louis Béthune +6

Diffusion (Large) Language Models (dLLMs) now match the downstream performance of their autoregressive counterparts on many tasks, while holding the promise of being more efficient…

cs.LG2026

Efficient and Uncertainty-Aware Diffusion Framework for Offline-to-Online Reinforcement Learning

Ha Manh Bui, Metod Jazbec, Eric Nalisnick +1

Offline-to-Online Reinforcement Learning (O2O-RL) leverages an offline, pre-trained policy to minimize costly online interactions. Although data-efficient, O2O-RL is susceptible to…

cs.AI2026

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting

Andrea Wynn, Metod Jazbec, Charith Peris +4

Large language models (LLMs) can be influenced by harmful or irrelevant context, which can significantly harm model performance on downstream tasks. This motivates principled desig…