Learning a Single Neuron with Gradient Methods
arXiv:2001.05205
Abstract
We consider the fundamental problem of learning a single neuron using standard gradient methods. As opposed to previous works, which considered specific (and not always realistic) input distributions and activation functions , we ask whether a more general result is attainable, under milder assumptions. On the one hand, we show that some assumptions on the distribution and the activation function are necessary. On the other hand, we prove positive guarantees under mild assumptions, which go beyond those studied in the literature so far. We also point out and study the challenges in further strengthening and generalizing our results.
Fixed a small bug in the proof of Theorem 4.2
References in corpus (4)
Cited by in corpus (8)
- Statistical-Query Lower Bounds via Functional Gradients
- Early-stopped neural networks are consistent
- Implicit Regularization in ReLU Networks with the Square Loss
- The Connection Between Approximation, Depth Separation and Learnability in Neural Networks
- The Effects of Mild Over-parameterization on the Optimization Landscape of Shallow ReLU Neural Networks
- Understanding How Over-Parametrization Leads to Acceleration: A case of learning a single teacher neuron
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear Network
- Directional Convergence Analysis under Spherically Symmetric Distribution