3 papers
cs.LG2026
Load Testing for Machine Learning Model Serving Systems at Scale
Amr S. Abdelfattah, Nakul Tirumalai, Indu Mohanan +4
Machine learning (ML) model serving has become a dominant consumer of GPU infrastructure, yet capacity planning in these systems remains largely ad hoc. Under-provisioning leads to…
cs.LG2025
StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
Qijun Luo, Mengqi Li, Lei Zhao +1
Training language models on long sequence data is a demanding requirement for enhancing the model's capability on complex tasks, e.g., long-chain reasoning. However, as the sequenc…
cs.LG2024
BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language Models
Qijun Luo, Hengxu Yu, Xiao Li
This work presents BAdam, an optimization method that leverages the block coordinate descent (BCD) framework with Adam's update rule. BAdam offers a memory efficient approach to th…