198 citations · 207 across the 4 of their papers we have counts for
4 papers
Generalisation dynamics of online learning in over-parameterised neural networks
Sebastian Goldt, Madhu S. Advani, Andrew M. Saxe +2
Deep neural networks achieve stellar generalisation on a variety of problems, despite often being large enough to easily fit all their training data. Here we study the generalisati…
Hierarchy through Composition with Linearly Solvable Markov Decision Processes
Andrew M. Saxe, Adam Earle, Benjamin Rosman
Hierarchical architectures are critical to the scalability of reinforcement learning methods. Current hierarchical frameworks execute actions serially, with macro-actions comprisin…
Tensor Switching Networks
Chuan-Yung Tsai, Andrew Saxe, David Cox
We present a novel neural network algorithm, the Tensor Switching (TS) network, which generalizes the Rectified Linear Unit (ReLU) nonlinearity to tensor-valued hidden units. The T…
Qualitatively characterizing neural network optimization problems
Ian J. Goodfellow, Oriol Vinyals, Andrew M. Saxe
Training neural networks involves solving large-scale non-convex optimization problems. This task has long been believed to be extremely difficult, with fear of local minima and ot…