1 paper · 1 filter
Neil Mallinar, Daniel Beaglehole, Libin Zhu +3
Neural networks trained to solve modular arithmetic tasks exhibit grokking, a phenomenon where the test accuracy starts improving long after the model achieves 100% training accura…