Showing stat.MLShow all
2 papers · 1 filter
stat.ML2025
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
Neil Mallinar, Daniel Beaglehole, Libin Zhu +3
Neural networks trained to solve modular arithmetic tasks exhibit grokking, a phenomenon where the test accuracy starts improving long after the model achieves 100% training accura…
stat.ML2024
Fast training of large kernel models with delayed projections
Amirhesam Abedsoltan, Siyuan Ma, Parthe Pandit +1
Classical kernel machines have historically faced significant challenges in scaling to large datasets and model sizes--a key ingredient that has driven the success of neural networ…