3 papers
cs.CY2020
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart +4
We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attai…
cs.CV2019
DIODE: A Dense Indoor and Outdoor DEpth Dataset
Igor Vasiljevic, Nick Kolkin, Shanyi Zhang +8
We introduce DIODE, a dataset that contains thousands of diverse high resolution color images with accurate, dense, long-range depth measurements. DIODE (Dense Indoor/Outdoor DEpth…
cs.LG2019
Natural Adversarial Examples
Dan Hendrycks, Kevin Zhao, Steven Basart +2
We introduce two challenging datasets that reliably cause machine learning model performance to substantially degrade. The datasets are collected with a simple adversarial filtrati…