From the 1 of 6 linked papers with an AI index.
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
Yuning Zhang, Grant Pinkert, Nan Yang +2
Large Language Models (LLMs) are increasingly deployed as Internet/Web services (LLM-as-a-Service) with strict latency Service-Level Objectives (SLOs) under tight GPU memory budget…
cs.DC2024
Threats and Defenses in Federated Learning Life Cycle: A Comprehensive Survey and Challenges
Yanli Li, Zhongliang Guo, Nan Yang +3
Federated Learning (FL) offers innovative solutions for privacy-preserving collaborative machine learning (ML). Despite its promising potential, FL is vulnerable to various attacks…