ADADELTA: An Adaptive Learning Rate Method
arXiv:1212.5701
Abstract
We present a novel per-dimension learning rate method for gradient descent called ADADELTA. The method dynamically adapts over time using only first order information and has minimal computational overhead beyond vanilla stochastic gradient descent. The method requires no manual tuning of a learning rate and appears robust to noisy gradient information, different model architecture choices, various data modalities and selection of hyperparameters. We show promising results compared to other methods on the MNIST digit classification task using a single machine and on a large scale voice dataset in a distributed cluster environment.
6 pages
Cited by in corpus (106)
- Deep Learning in Neural Networks: An Overview
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Attention-Based Models for Speech Recognition
- Convolutional Neural Networks for Sentence Classification
- Dual Learning for Machine Translation
- On Using Monolingual Corpora in Neural Machine Translation
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Scalable Variational Gaussian Process Classification
- Gravity Spy: Integrating Advanced LIGO Detector Characterization, Machine Learning, and Citizen Science
- MADE: Masked Autoencoder for Distribution Estimation
- Text Classification Improved by Integrating Bidirectional LSTM with Two-dimensional Max Pooling
- Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
- Toward Multilingual Neural Machine Translation with Universal Encoder and Decoder
- RETAIN: An Interpretable Predictive Model for Healthcare using Reverse Time Attention Mechanism
- Learning Stochastic Recurrent Networks
- Is Neural Machine Translation Ready for Deployment? A Case Study on 30 Translation Directions
- Discovering Hidden Factors of Variation in Deep Networks
- Protein Secondary Structure Prediction with Long Short Term Memory Networks
- Deep Learning for Medical Image Segmentation
- Character-level and Multi-channel Convolutional Neural Networks for Large-scale Authorship Attribution
- On the Origin of Deep Learning
- Two are Better than One: An Ensemble of Retrieval- and Generation-Based Dialog Systems
- Classify or Select: Neural Architectures for Extractive Document Summarization
- Hierarchical Attention Network for Action Recognition in Videos
- Applications of Online Deep Learning for Crisis Response Using Social Media Information
- Collaborative Filtering with Recurrent Neural Networks
- Chinese Poetry Generation with Planning based Neural Network
- Attending to Characters in Neural Sequence Labeling Models
- Distraction-Based Neural Networks for Document Summarization
- Zero-Shot Visual Question Answering
- Learning What Data to Learn
- Temporal Attention Model for Neural Machine Translation
- Deep learning for plasma tomography using the bolometer system at JET
- Time Series Classification from Scratch with Deep Neural Networks: A Strong Baseline
- Neural Sentence Ordering
- Alzheimer's Disease Diagnostics by a Deeply Supervised Adaptable 3D Convolutional Network
- Discriminative Acoustic Word Embeddings: Recurrent Neural Network-Based Approaches
- Multi-task Prediction of Disease Onsets from Longitudinal Lab Tests
- Rapid Classification of Crisis-Related Data on Social Networks using Convolutional Neural Networks
- Exploiting Convolutional Neural Network for Risk Prediction with Medical Feature Embedding
- Recurrent Neural Networks with External Memory for Language Understanding
- Doubly Convolutional Neural Networks
- Latent Variable Dialogue Models and their Diversity
- Word Embeddings and Their Use In Sentence Classification Tasks
- Supervised Attentions for Neural Machine Translation
- Neural Machine Translation with Supervised Attention
- Neural Emoji Recommendation in Dialogue Systems
- Answer Sequence Learning with Neural Networks for Answer Selection in Community Question Answering
- Interactive Attention for Neural Machine Translation
- Attention-Based Multimodal Fusion for Video Description
- One Sentence One Model for Neural Machine Translation
- Iterative Neural Autoregressive Distribution Estimator (NADE-k)
- Comparing Rule-Based and Deep Learning Models for Patient Phenotyping
- Deep Robust Kalman Filter
- Joint CTC-Attention based End-to-End Speech Recognition using Multi-task Learning
- Generative Adversarial Networks as Variational Training of Energy Based Models
- Character-level Convolutional Network for Text Classification Applied to Chinese Corpus
- Improving Interpretability of Deep Neural Networks with Semantic Information
- Deep Recurrent Models with Fast-Forward Connections for Neural Machine Translation
- Neural Networks Models for Entity Discovery and Linking
- Latent Gaussian Processes for Distribution Estimation of Multivariate Categorical Data
- Activation Ensembles for Deep Neural Networks
- Nematus: a Toolkit for Neural Machine Translation
- Denoising autoencoder with modulated lateral connections learns invariant representations of natural images
- Trainable Greedy Decoding for Neural Machine Translation
- First Steps Toward Incorporating Image Based Diagnostics Into Particle Accelerator Control Systems Using Convolutional Neural Networks
- Prediction of Kidney Function from Biopsy Images Using Convolutional Neural Networks
- Using Sentence Plausibility to Learn the Semantics of Transitive Verbs
- SurvivalNet: Predicting patient survival from diffusion weighted magnetic resonance images using cascaded fully convolutional and 3D convolutional neural networks
- Span-Based Constituency Parsing with a Structure-Label System and Provably Optimal Dynamic Oracles
- An Empirical Exploration of Skip Connections for Sequential Tagging
- Recurrent Neural Network based Part-of-Speech Tagger for Code-Mixed Social Media Text
- Neural Machine Translation Advised by Statistical Machine Translation
- Trusting SVM for Piecewise Linear CNNs
- Multimodal Memory Modelling for Video Captioning
- Generative Knowledge Transfer for Neural Language Models
- Character-Level Language Modeling with Hierarchical Recurrent Neural Networks
- Scalable Gaussian Process Classification via Expectation Propagation
- Word and Document Embeddings based on Neural Network Approaches
- Smart Library: Identifying Books in a Library using Richly Supervised Deep Scene Text Reading
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- Dependency Sensitive Convolutional Neural Networks for Modeling Sentences and Documents
- Learning Semantically Coherent and Reusable Kernels in Convolution Neural Nets for Sentence Classification
- FPGA-Based Low-Power Speech Recognition with Recurrent Neural Networks
- An Efficient Character-Level Neural Machine Translation
- Learning From Graph Neighborhoods Using LSTMs
- Hot Swapping for Online Adaptation of Optimization Hyperparameters
- Efficient Elastic Net Regularization for Sparse Linear Models
- Autoencoder Regularized Network For Driving Style Representation Learning
- Neural-based Noise Filtering from Word Embeddings
- The AMU-UEDIN Submission to the WMT16 News Translation Task: Attention-based NMT Models as Feature Functions in Phrase-based SMT
- S3Pool: Pooling with Stochastic Spatial Sampling
- Siamese convolutional networks based on phonetic features for cognate identification
- Unsupervised preprocessing for Tactile Data
- Hybrid Dialog State Tracker with ASR Features
- Towards Deep Compositional Networks
- Distinguishing Antonyms and Synonyms in a Pattern-based Neural Network
- Neural Multi-Source Morphological Reinflection
- A Machine-Learning Framework for Design for Manufacturability
- Toward Implicit Sample Noise Modeling: Deviation-driven Matrix Factorization
- On SGD's Failure in Practice: Characterizing and Overcoming Stalling
- Charged Point Normalization: An Efficient Solution to the Saddle Point Problem
- Single image super-resolution using self-optimizing mask via fractional-order gradient interpolation and reconstruction
- Representation Learning Models for Entity Search
- Greedy Step Averaging: A parameter-free stochastic optimization method
- Neural Networks Classifier for Data Selection in Statistical Machine Translation