10 papers
LMEdge: QoS-Aware LLM Inference Orchestration on Edge Clusters
Reza Farahani, Zoha Azimi, Mario Colosi +1
Large language model (LLM) services increasingly operate on edge infrastructure, enabling low-latency and privacy-preserving AI services. However, efficiently serving LLM requests…
Scale: Deep Reinforcement Learning for Container Scheduling in Serverless Edge Computing
Chen Chen, Zihan Jia, Andrea Sabbioni +2
Serverless computing has emerged as a promising computing paradigm for edge computing. However, adopting the event driven model in highly dynamic, heterogeneous, and distributed ed…
Orchestrating Serverless Applications in the Edge Cloud Space Continuum: What Breaks and What is Next?
Hadi Tabatabaee Malazi, Reza Farahani, Nitinder Mohan +1
Serverless computing has matured into an effective execution model for edge cloud environments, enabling function level decomposition, demand driven scaling, and workflow execution…
ClusterLess: Deadline-Aware Serverless Workflow Orchestration on Federated Edge Clusters
Reza Farahani, Mario Colosi, Ilir Murturi +4
The recent convergence of edge computing, serverless execution, and Kubernetes (K8s) based container orchestration has enabled the processing of application workflows close to data…
DQ-Ladder: A Deep Reinforcement Learning-based Bitrate Ladder for Adaptive Video Streaming
Reza Farahani, Zoha Azimi, Vignesh V Menon +4
Adaptive streaming of segmented video over HTTP typically relies on a predefined set of bitrate-resolution pairs, known as a bitrate ladder. However, fixed ladders often overlook v…
Real-Time AI Service Economy: A Framework for Agentic Computing Across the Continuum
Lauri Lovén, Alaa Saleh, Reza Farahani +4
Real-time AI services increasingly operate across the device-edge-cloud continuum, where autonomous AI agents generate latency-sensitive workloads, orchestrate multi-stage processi…