25 citations · 25 across the 12 of their papers we have counts for
14 papers
RouteBalance: Fused Model Routing and Load Balancing for Heterogeneous LLM Serving
Wei Da, Evangelia Kalyvianaki
Heterogeneous LLM serving stacks split scheduling into two layers that optimize in isolation: model routers pick a model from quality and cost signals while ignoring instance load,…
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
Wei Da, Evangelia Kalyvianaki
Realistic evaluation of LLM serving systems requires online workloads, dynamic arrivals, queueing, and the serving engine's local scheduling for execution batching, but running suc…
Dodoor: Efficient Randomized Decentralized Scheduling with Load Caching for Heterogeneous Tasks and Clusters
Wei Da, Evangelia Kalyvianaki
This paper presents Dodoor, a randomized decentralized scheduler for heterogeneous clusters. Dodoor removes hot-path probing via batched cache refreshes and introduces a heterogene…
Mitigating context switching in densely packed Linux clusters with Latency-Aware Group Scheduling
Al Amjad Tawfiq Isstaif, Evangelia Kalyvianaki, Richard Mortier
Cluster orchestrators such as Kubernetes depend on accurate estimates of node capacity and job requirements. Inaccuracies in either lead to poor placement decisions and degraded cl…
Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling
Wei Da, Evangelia Kalyvianaki
This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language model (LLM) serving. Astrolabe improves load…
Average Consensus over Directed Networks in Open Multi-Agent Systems with Acknowledgement Feedback
Evagoras Makridis, Andreas Grammenos, Gabriele Oliva +3
In this paper, we address the distributed average consensus problem over directed networks in open multi-agent systems (OMAS), where the stability of the network is disrupted by fr…