10 citations · 10 across the 3 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.LG2022
Neural Attentive Circuits
Nasim Rahaman, Martin Weiss, Francesco Locatello +5
Recent work has seen the development of general purpose neural architectures that can be trained to perform tasks across diverse data modalities. General purpose models typically m…
cs.LG2022★ 10 cited
Non-Convergence and Limit Cycles in the Adam optimizer
Sebastian Bock, Martin Georg Weiß
One of the most popular training algorithms for deep neural networks is the Adaptive Moment Estimation (Adam) introduced by Kingma and Ba. Despite its success in many applications…