3 papers
cs.LG2026
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
Hiroki Naganuma, Shagun Gupta, Youssef Briki +4
To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, relying on heuristics that are brittle and…
cs.LG2026
Takeuchi's Information Criteria as Generalization Measures for DNNs Close to NTK Regime
Hiroki Naganuma, Taiji Suzuki, Rio Yokota +3
Generalization measures have been studied extensively in the machine learning community to better characterize generalization gaps. However, establishing a reliable generalization…
cs.LG2025
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
Hiroki Naganuma, Ryuichiro Hataya, Kotaro Yoshida +1
In the field of computer vision, fine-tuning pre-trained models has become a prevalent strategy for out-of-distribution (OOD) generalization tasks. Different from most prior work t…