402 citations · 404 across the 2 of their papers we have counts for
7 papers
How Would The Viewer Feel? Estimating Wellbeing From Video Scenarios
Mantas Mazeika, Eric Tang, Andy Zou +6
In recent years, deep neural networks have demonstrated increasingly strong abilities to recognize objects and activities in videos. However, as video understanding becomes widely…
Measuring Coding Challenge Competence With APPS
Dan Hendrycks, Steven Basart, Saurav Kadavath +8
While programming is one of the most broadly applicable skills in modern society, modern machine learning models still cannot code solutions to basic problems. Despite its importan…
Measuring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart +4
We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attai…
Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty
Dan Hendrycks, Mantas Mazeika, Saurav Kadavath +1
Self-supervision provides effective representations for downstream tasks without requiring labels. However, existing approaches lag behind fully supervised training and are often n…
Deep Anomaly Detection with Outlier Exposure
Dan Hendrycks, Mantas Mazeika, Thomas Dietterich
It is important to detect anomalous inputs when deploying machine learning systems. The use of larger and more complex inputs in deep learning magnifies the difficulty of distingui…
Using Pre-Training Can Improve Model Robustness and Uncertainty
Dan Hendrycks, Kimin Lee, Mantas Mazeika
He et al. (2018) have called into question the utility of pre-training by showing that training from scratch can often yield similar performance to pre-training. We show that altho…