Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Online Safety Monitoring for LLMs
Mona Schirmer, Metod Jazbec, Alexander Timans +3
Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed i…
cs.AI2026
Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting
Andrea Wynn, Metod Jazbec, Charith Peris +4
Large language models (LLMs) can be influenced by harmful or irrelevant context, which can significantly harm model performance on downstream tasks. This motivates principled desig…