9 papers · 1 filter
LiMuon: Light and Fast Muon Optimizer for Large Models
Feihu Huang, Yuning Luo, Songcan Chen
Large models recently are widely applied in machine learning, so efficient training of large models has received widespread attention. More recently, the useful Muon optimizer is s…
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
Feihu Huang, Yuning Luo, Songcan Chen
Matrix-structured parameters frequently appear in many artificial intelligence models such as large language models. More recently, an efficient Muon optimizer is designed for matr…
CLion: Efficient Cautious Lion Optimizer with Enhanced Generalization
Feihu Huang, Guanyi Zhang, Songcan Chen
Lion optimizer is a popular learning-based optimization algorithm in machine learning, which shows impressive performance in training many deep learning models. Although convergenc…
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
Feihu Huang, Guanyi Zhang, Songcan Chen
Adam and AdamW are a class of default optimizers for training deep learning models in machine learning. These adaptive algorithms converge faster but generalize worse compared to S…
Negatives-Dominant Contrastive Learning for Generalization in Imbalanced Domains
Meng Cao, Jiexi Liu, Songcan Chen
Imbalanced Domain Generalization (IDG) focuses on mitigating both domain and label shifts, both of which fundamentally shape the model's decision boundaries, particularly under het…
Beyond Observations: Reconstruction Error-Guided Irregularly Sampled Time Series Representation Learning
Jiexi Liu, Meng Cao, Songcan Chen
Irregularly sampled time series (ISTS), characterized by non-uniform time intervals with natural missingness, are prevalent in real-world applications. Existing approaches for ISTS…