5 papers
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
Hiroki Naganuma, Shagun Gupta, Youssef Briki +4
To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, relying on heuristics that are brittle and…
Navigating Potholes with Geometry-Aware Sharpness Minimization
Simon Dufort-Labbé, Mehrab Hamidi, Razvan Pascanu +3
Sharpness-aware minimization (SAM) encourages flat minima by perturbing parameters along directions of high loss curvature, but treats all parameter directions uniformly, ignoring…
Generalization Measures under Controlled Covariate Shift: A Regime-Aware Benchmark
Sora Nakai, Youssef Fadhloun, Kacem Mathlouthi +4
Predicting generalization from quantities available before target-test evaluation remains a central challenge in deep learning. The systematic benchmark of Jiang et al. (2020) eval…
Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training
Hiroki Naganuma, Xinzhi Zhang, Man-Chung Yue +4
Following AI scaling trends, frontier models continue to grow in size and continue to be trained on larger datasets. Training these models requires huge investments in exascale com…
Solving Hidden Monotone Variational Inequalities with Surrogate Losses
Ryan D'Orazio, Danilo Vucetic, Zichu Liu +3
Deep learning has proven to be effective in a wide variety of loss minimization problems. However, many applications of interest, like minimizing projected Bellman error and min-ma…