1k citations · 2.4k across the 12 of their papers we have counts for
18 papers · 1 filter
New CleverHans Feature: Better Adversarial Robustness Evaluations with Attack Bundling
Ian Goodfellow
This technical report describes a new feature of the CleverHans library called "attack bundling". Many papers about adversarial examples present lists of error rates corresponding…
Local Explanation Methods for Deep Neural Networks Lack Sensitivity to Parameter Values
Julius Adebayo, Justin Gilmer, Ian Goodfellow +1
Explaining the output of a complicated machine learning model like a deep neural network (DNN) is a central challenge in machine learning. Several proposed local explanation method…
Discriminator Rejection Sampling
Samaneh Azadi, Catherine Olsson, Trevor Darrell +2
We propose a rejection sampling scheme using the discriminator of a GAN to approximately correct errors in the GAN generator distribution. We show that under quite strict assumptio…
Sanity Checks for Saliency Maps
Julius Adebayo, Justin Gilmer, Michael Muelly +3
Saliency methods have emerged as a popular tool to highlight features in an input deemed relevant for the prediction of a learned model. Several saliency methods have been proposed…
Unrestricted Adversarial Examples
Tom B. Brown, Nicholas Carlini, Chiyuan Zhang +3
We introduce a two-player contest for evaluating the safety and robustness of machine learning systems, with a large prize pool. Unlike most prior work in ML robustness, which stud…
Skill Rating for Generative Models
Catherine Olsson, Surya Bhupatiraju, Tom Brown +2
We explore a new way to evaluate generative models using insights from evaluation of competitive games between human players. We show experimentally that tournaments between genera…