Showing 2025Show all
3 papers · 1 filter
cs.LG2025
FOAM: Blocked State Folding for Memory-Efficient LLM Training
Ziqing Wen, Jiahuan Wang, Ping Luo +2
Large language models (LLMs) have demonstrated remarkable performance due to their large parameter counts and extensive training data. However, their scale leads to significant mem…
cs.AI2025
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
Mengkang Hu, Yuhang Zhou, Wendong Fan +13
Large Language Model (LLM)-based multi-agent systems show promise for automating real-world tasks but struggle to transfer across domains due to their domain-specific nature. Curre…
cs.LG2025
GWT: Scalable Optimizer State Compression for Large Language Model Training
Ziqing Wen, Ping Luo, Jiahuan Wang +4
Training large language models (LLMs) requires substantial memory, a significant fraction of which is consumed by the moment states maintained by adaptive optimizers such as Adam.…