2 papers
cs.LG2025
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
Qingao Yi, Jiaang Duan, Hanwen Hu +10
Training large language models (LLMs) poses significant challenges regarding computational resources and memory capacity. Although distributed training techniques help mitigate the…
cs.DC2025
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
Jiaang Duan, Shenglin Xu, Shiyou Qian +15
The surge in large language models (LLMs) has fundamentally reshaped the landscape of GPU usage patterns, creating an urgent need for more efficient management strategies. While cl…