1 citations · 1 across the 2 of their papers we have counts for
3 papers
Detecting and Controlling Sycophancy with Cascading Linear Features
Maty Bohacek, Rishub Jain, Nicholas Dufour +3
Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desired or undesired behavior. Thes…
Can AI mediation improve democratic deliberation?
Michael Henry Tessler, Georgina Evans, Michiel A. Bakker +8
The strength of democracy lies in the free and equal exchange of diverse viewpoints. Living up to this ideal at scale faces inherent tensions: broad participation, meaningful delib…
An Approach to Technical AGI Safety and Security
Rohin Shah, Alex Irpan, Alexander Matt Turner +27
Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough…