1 paper · 1 filter
Wesley Hanwen Deng, Sunnie S. Y. Kim, Akshita Jha +4
Recent developments in AI governance and safety research have called for red-teaming methods that can effectively surface potential risks posed by AI models. Many of these calls ha…