2 papers
cs.LG2025
Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking
Athanasios Glentis, Jiaxiang Li, Qiulin Shang +4
Fueled by their remarkable ability to tackle diverse tasks across multiple domains, large language models (LLMs) have grown at an unprecedented rate, with some recent models contai…
math.OC2024
A Mathematics-Inspired Learning-to-Optimize Framework for Decentralized Optimization
Yutong He, Qiulin Shang, Xinmeng Huang +2
Most decentralized optimization algorithms are handcrafted. While endowed with strong theoretical guarantees, these algorithms generally target a broad class of problems, thereby n…