3 papers
cs.DC2026
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
Rui Li, Zhaoning Zhang, Libo Zhang +3
Speculative decoding (SD) accelerates LLM inference by verifying draft tokens in parallel. However, this method presents a critical trade-off: it improves throughput in low-load, m…
cs.DC2026
Joint: Orchestrating Serverless Workflows on Jointcloud FaaS Systems
Rui Li, Jianfei Liu, Zhilin Yang +3
Existing serverless workflow orchestration systems are predominantly designed for a single-cloud FaaS system, leading to vendor lock-in. This restricts performance optimization, co…
cs.DC2024
NebulaFL: Effective Asynchronous Federated Learning for JointCloud Computing
Fei Gao, Ming Hu, Zhiyu Xie +4
With advancements in AI infrastructure and Trusted Execution Environment (TEE) technology, Federated Learning as a Service (FLaaS) through JointCloud Computing (JCC) is promising t…