2 papers
cs.AI2026
Online Safety Monitoring for LLMs
Mona Schirmer, Metod Jazbec, Alexander Timans +3
Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed i…
stat.ML2026
CAOS: Conformal Aggregation of One-Shot Predictors
Maja Waldron
One-shot prediction enables rapid adaptation of pretrained foundation models to new tasks using only one labeled example, but lacks principled uncertainty quantification. While con…