Identifying and Exploiting Structures for Reliable Deep Learning
arXiv:2108.07083
Abstract
Deep learning research has recently witnessed an impressively fast-paced progress in a wide range of tasks including computer vision, natural language processing, and reinforcement learning. The extraordinary performance of these systems often gives the impression that they can be used to revolutionise our lives for the better. However, as recent works point out, these systems suffer from several issues that make them unreliable for use in the real world, including vulnerability to adversarial attacks (Szegedy et al. [248]), tendency to memorise noise (Zhang et al. [292]), being over-confident on incorrect predictions (miscalibration) (Guo et al. [99]), and unsuitability for handling private data (Gilad-Bachrach et al. [88]). In this thesis, we look at each of these issues in detail, investigate their causes, and propose computationally cheap algorithms for mitigating them in practice. To do this, we identify structures in deep neural networks that can be exploited to mitigate the above causes of unreliability of deep learning algorithms.
Final Thesis for DPhil in Computer Science submitted at the University of Oxford
References in corpus (25)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- Conditional Generative Adversarial Nets
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Understanding deep learning requires rethinking generalization
- A Learned Representation For Artistic Style
- Certified Adversarial Robustness via Randomized Smoothing
- Compressing Neural Networks with the Hashing Trick
- Fast is better than free: Revisiting adversarial training
- Private federated learning on vertically partitioned data via entity resolution and additively homomorphic encryption
- Regularizing Neural Networks by Penalizing Confident Output Distributions
- Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients
- Spectral Norm Regularization for Improving the Generalizability of Deep Learning
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Beyond temperature scaling: Obtaining well-calibrated multiclass probabilities with Dirichlet calibration
- Improving Neural Language Models with a Continuous Cache
- Evaluating model calibration in classification
- Complexity of Linear Regions in Deep Networks
- Deterministic PAC-Bayesian generalization bounds for deep networks via generalizing noise-resilience
- Towards GAN Benchmarks Which Require Generalization
- Diverse Ensembles Improve Calibration
- Multiscale sequence modeling with a learned dictionary
- DeepSecure: Scalable Provably-Secure Deep Learning