2 papers
cs.DC2025
MoFa: A Unified Performance Modeling Framework for LLM Pretraining
Lu Zhao, Rong Shi, Shaoqing Zhang +14
The exponential growth in LLM scales, with parameters soaring from billions to trillions, has necessitated distributed pretraining across large clusters comprising thousands to ten…
cs.DC2025
FedCostAware: Enabling Cost-Aware Federated Learning on the Cloud
Aditya Sinha, Zilinghan Li, Tingkai Liu +3
Federated learning (FL) is a distributed machine learning (ML) approach that allows multiple clients to collaboratively train ML models without exchanging original training data, o…