2 papers
cs.LG2026
Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio
Ziqing Wen, Zhouyang Liu, Jiahuan Wang +4
The impressive performance of large language models (LLMs) arises from their massive scale and heterogeneous module composition. However, this structural heterogeneity introduces a…
cs.LG2024
Exploring the Generalization Capabilities of AID-based Bi-level Optimization
Congliang Chen, Li Shen, Zhiqiang Xu +3
Bi-level optimization has achieved considerable success in contemporary machine learning applications, especially for given proper hyperparameters. However, due to the two-level op…