1 citations · 1 across the 2 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Detecting and Controlling Sycophancy with Cascading Linear Features
Maty Bohacek, Rishub Jain, Nicholas Dufour +3
Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desired or undesired behavior. Thes…
cs.AI2025★ 1 cited
An Approach to Technical AGI Safety and Security
Rohin Shah, Alex Irpan, Alexander Matt Turner +27
Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough…