Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
Smoothness Adaptivity in Constant-Depth Neural Networks: Optimal Rates via Smooth Activations
Yuhao Liu, Zilin Wang, Lei Wu +1
Smooth activation functions are ubiquitous in modern deep learning, yet their theoretical advantages over non-smooth counterparts remain poorly understood. In this work, we study b…
stat.ML2026
Optimal Learning-Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay
Binghui Li, Zilin Wang, Fengling Chen +3
We study optimal learning-rate schedules (LRSs) under the functional scaling law (FSL) framework introduced in Li et al. (2025), which accurately models the loss dynamics of both l…