most citedInvestigating the interaction between gradient-only line searches and different activation functions

1 citations · 2 across the 3 of their papers we have counts for

collaborators

6 papers

stat.ML20201 cited

Gradient-only line searches to automatically determine learning rates for a variety of stochastic training algorithms

Dominic Kafka, Daniel Nicolas Wilke

Gradient-only and probabilistic line searches have recently reintroduced the ability to adaptively determine learning rates in dynamic mini-batch sub-sampled neural network trainin…

stat.ML20201 cited

Investigating the interaction between gradient-only line searches and different activation functions

D. Kafka, Daniel. N. Wilke

Gradient-only line searches (GOLS) adaptively determine step sizes along search directions for discontinuous loss functions resulting from dynamic mini-batch sub-sampling in neural…

stat.ML2020

Resolving learning rates adaptively by locating Stochastic Non-Negative Associated Gradient Projection Points using line searches

Dominic Kafka, Daniel N. Wilke

Learning rates in stochastic neural network training are currently determined a priori to training, using expensive manual or automated iterative tuning. This study proposes gradie…

stat.ML2019

Empirical study towards understanding line search approximations for training neural networks

Younghwan Chae, Daniel N. Wilke

Choosing appropriate step sizes is critical for reducing the computational cost of training large-scale neural network models. Mini-batch sub-sampling (MBSS) is often employed for…

stat.ML2019

Gradient-only line searches: An Alternative to Probabilistic Line Searches

Dominic Kafka, Daniel Wilke

Step sizes in neural network training are largely determined using predetermined rules such as fixed learning rates and learning rate schedules. These require user input or expensi…

stat.ML2019

Traversing the noise of dynamic mini-batch sub-sampled loss functions: A visual guide

Dominic Kafka, Daniel Wilke

Mini-batch sub-sampling in neural network training is unavoidable, due to growing data demands, memory-limited computational resources such as graphical processing units (GPUs), an…