Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
Dylan Feng, Pragya Srivastava, Anca Dragan +1
Many safety and alignment failures of large language models (LLMs) occur due to out-of-distribution (OOD) situations: unusual prompt or response patterns that are unforeseen by mod…
cs.AI2025
AssistanceZero: Scalably Solving Assistance Games
Cassidy Laidlaw, Eli Bronstein, Timothy Guo +5
Assistance games are a promising alternative to reinforcement learning from human feedback (RLHF) for training AI assistants. Assistance games resolve key drawbacks of RLHF, such a…