260 citations · 331 across the 15 of their papers we have counts for
22 papers · 1 filter
Coordinated Scheduling for MoE LLM Serving
Yifan Sun, Zhexiang Zhang, Jiantong Jiang +5
Serving Mixture-of-Experts (MoE) large language models (LLMs) is challenging because dynamic request workloads interact with sparse expert routing, creating both data-parallel (DP)…
Multi-Layer Scheduling for MoE-Based LLM Reasoning
Yifan Sun, Gholamreza Haffari, Minxian Xu +2
Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks, but serving them efficiently at scale remains a critical challenge due to their substant…
LLM-Driven Intent-Based Privacy-Aware Orchestration Across the Cloud-Edge Continuum
Zijie Su, Muhammed Tawfiqul Islam, Mohammad Goudarzi +1
With the rapid advancement of large language models (LLMs), efficiently serving LLM inference under limited GPU resources has become a critical challenge. Recently, an increasing n…
DistributedEstimator: Distributed Training of Quantum Neural Networks via Circuit Cutting
Prabhjot Singh, Adel N. Toosi, Rajkumar Buyya
Circuit cutting decomposes a large quantum circuit into smaller subcircuits executed independently; expectation values are recovered by classically combining subcircuit outcomes. P…
REACH: Reinforcement Learning for Adaptive Microservice Rescheduling in the Cloud-Edge Continuum
Xu Bai, Muhammed Tawfiqul Islam, Rajkumar Buyya +1
Cloud computing, despite its advantages in scalability, may not always fully satisfy the low-latency demands of emerging latency-sensitive pervasive applications. The cloud-edge co…
Resilience Evaluation of Kubernetes in Cloud-Edge Environments via Failure Injection
Zihao Chen, Mohammad Goudarzi, Adel Nadjaran Toosi
Kubernetes has emerged as an essential platform for deploying containerised applications across cloud and edge infrastructures. As Kubernetes gains increasing adoption for mission-…