2 papers
cs.AI2026
Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption
Mert Yildiz, Pietro Spadaccino, Alexey Rolich +2
Modern deployments of Large Language Models (LLMs) increasingly require serving multiple models with diverse architectures, sizes, and specialization on shared, heterogeneous hardw…
cs.NI2025
SPARQ: An Optimization Framework for the Distribution of AI-Intensive Applications under Non-Linear Delay Constraints
Pietro Spadaccino, Paolo Di Lorenzo, Sergio Barbarossa +2
Next-generation real-time compute-intensive applications, such as extended reality, multi-user gaming, and autonomous transportation, are increasingly composed of heterogeneous AI-…