Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Human-AI Complementarity: A Goal for Amplified Oversight
Rishub Jain, Sophie Bridgers, Lili Janzer +3
Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging tasks, verifying quality and safety becomes…
cs.AI2026
Detecting and Controlling Sycophancy with Cascading Linear Features
Maty Bohacek, Rishub Jain, Nicholas Dufour +3
Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desired or undesired behavior. Thes…
cs.AI2025
An Approach to Technical AGI Safety and Security
Rohin Shah, Alex Irpan, Alexander Matt Turner +27
Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough…