2 papers
cs.LG2026
Amortizing Scaling Law Construction Costs
Abhash Kumar Jha, Diana Alexandra Onuţu, Neeratyoy Mallik +6
Scaling laws guide the design choices for training large foundation models, but deriving them involves training an exhaustive grid over hyperparameters, token budgets, and paramete…
cs.LG2025
Adam Simplified: Bias Correction Debunked
Sam Laing, Antonio Orvieto
The Adam optimizer is a cornerstone of modern deep learning, yet the empirical necessity of each of its individual components is often taken for granted. This paper presents a focu…