2 papers
cs.LG2026
Debate Training Reduces Reward Hacking in RLAIF
Zachary Kenton, Lili Janzer, Rory Greig +8
We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking comp…
cs.AI2025
Human-AI Complementarity: A Goal for Amplified Oversight
Rishub Jain, Sophie Bridgers, Lili Janzer +3
Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging tasks, verifying quality and safety becomes…