1 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 1 cited
Red Teaming Deep Neural Networks with Feature Synthesis Tools
Stephen Casper, Yuxiao Li, Jiawei Li +4
Interpretable AI tools are often motivated by the goal of understanding model behavior in out-of-distribution (OOD) contexts. Despite the attention this area of study receives, the…
cs.LG2022★ 1 cited
Diagnostics for Deep Neural Networks with Automated Copy/Paste Attacks
Stephen Casper, Kaivalya Hariharan, Dylan Hadfield-Menell
This paper considers the problem of helping humans exercise scalable oversight over deep neural networks (DNNs). Adversarial examples can be useful by helping to reveal weaknesses…