2 papers
stat.ML2026
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
Libin Zhu, Damek Davis, Dmitriy Drusvyatskiy +1
We study a prototypical situation when a learned predictor can discover useful low-dimensional structure in data, while using fewer samples than are needed for accurate prediction.…
stat.ML2025
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
Neil Mallinar, Daniel Beaglehole, Libin Zhu +3
Neural networks trained to solve modular arithmetic tasks exhibit grokking, a phenomenon where the test accuracy starts improving long after the model achieves 100% training accura…