Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers
Siyuan Li, Jiabao Pan, Yumou Liu +9
Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by compute, memory, tuning budget, and task diversity, yet the land…
cs.LG2025
Taming LLMs by Scaling Learning Rates with Gradient Grouping
Siyuan Li, Juanxi Tian, Zedong Wang +4
Training large language models (LLMs) poses challenges due to their massive scale and heterogeneous architectures. While adaptive optimizers like AdamW help address gradient variat…
cs.LG2024
A Survey on Mixup Augmentations and Beyond
Xin Jin, Hongyu Zhu, Siyuan Li +6
As Deep Neural Networks have achieved thrilling breakthroughs in the past decade, data augmentations have garnered increasing attention as regularization techniques when massive la…