collaborators

5 papers

cs.LG2026

Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent

Hiroki Naganuma, Shagun Gupta, Youssef Briki +4

To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, relying on heuristics that are brittle and…

cs.LG2026

Navigating Potholes with Geometry-Aware Sharpness Minimization

Simon Dufort-Labbé, Mehrab Hamidi, Razvan Pascanu +3

Sharpness-aware minimization (SAM) encourages flat minima by perturbing parameters along directions of high loss curvature, but treats all parameter directions uniformly, ignoring…

cs.LG2026

Generalization Measures under Controlled Covariate Shift: A Regime-Aware Benchmark

Sora Nakai, Youssef Fadhloun, Kacem Mathlouthi +4

Predicting generalization from quantities available before target-test evaluation remains a central challenge in deep learning. The systematic benchmark of Jiang et al. (2020) eval…

cs.LG2025

Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training

Hiroki Naganuma, Xinzhi Zhang, Man-Chung Yue +4

Following AI scaling trends, frontier models continue to grow in size and continue to be trained on larger datasets. Training these models requires huge investments in exascale com…

cs.LG2025

Solving Hidden Monotone Variational Inequalities with Surrogate Losses

Ryan D'Orazio, Danilo Vucetic, Zichu Liu +3

Deep learning has proven to be effective in a wide variety of loss minimization problems. However, many applications of interest, like minimizing projected Bellman error and min-ma…