1 paper
Srijan Tiwari, Aditya Chauhan, Manjot Singh
Why do neural networks memorize algorithmic training data long before they generalize? We present a geometric case study demonstrating that, on tasks where generalization requires…