5 papers
FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving
Shouwei Gao, Junqi Yin, Feiyi Wang +1
Production LLM serving must simultaneously deliver high throughput, low latency, and sufficient context capacity under non-stationary traffic and mixed request requirements. Data p…
OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC
Sahil Tyagi, Andrei Cozma, Olivera Kotevska +1
Federated Learning (FL) is critical for edge and High Performance Computing (HPC) where data is not centralized and privacy is crucial. We present OmniFed, a modular framework desi…
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
Yueming Yuan, Ahan Gupta, Jianping Li +3
Emerging expert-specialized Mixture-of-Experts (MoE) architectures, such as DeepSeek-MoE, deliver strong model quality through fine-grained expert segmentation and large top-k rout…
Distributed Cross-Channel Hierarchical Aggregation for Foundation Models
Aristeidis Tsaris, Isaac Lyngaas, John Lagregren +6
Vision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from their ability to aggregate images fr…
AI-coupled HPC Workflow Applications, Middleware and Performance
Wes Brewer, Ana Gainaru, Frédéric Suter +3
AI integration is revolutionizing the landscape of HPC simulations, enhancing the importance, use, and performance of AI-driven HPC workflows. This paper surveys the diverse and ra…