7 papers · 1 filter
LMEdge: QoS-Aware LLM Inference Orchestration on Edge Clusters
Reza Farahani, Zoha Azimi, Mario Colosi +1
Large language model (LLM) services increasingly operate on edge infrastructure, enabling low-latency and privacy-preserving AI services. However, efficiently serving LLM requests…
Scale: Deep Reinforcement Learning for Container Scheduling in Serverless Edge Computing
Chen Chen, Zihan Jia, Andrea Sabbioni +2
Serverless computing has emerged as a promising computing paradigm for edge computing. However, adopting the event driven model in highly dynamic, heterogeneous, and distributed ed…
Orchestrating Serverless Applications in the Edge Cloud Space Continuum: What Breaks and What is Next?
Hadi Tabatabaee Malazi, Reza Farahani, Nitinder Mohan +1
Serverless computing has matured into an effective execution model for edge cloud environments, enabling function level decomposition, demand driven scaling, and workflow execution…
ClusterLess: Deadline-Aware Serverless Workflow Orchestration on Federated Edge Clusters
Reza Farahani, Mario Colosi, Ilir Murturi +4
The recent convergence of edge computing, serverless execution, and Kubernetes (K8s) based container orchestration has enabled the processing of application workflows close to data…
Serverless Everywhere: A Comparative Analysis of WebAssembly Workflows Across Browser, Edge, and Cloud
Mario Colosi, Reza Farahani, Lauri Loven +2
WebAssembly (Wasm) is a binary instruction format that enables portable, sandboxed, and near-native execution across heterogeneous platforms, making it well-suited for serverless w…
Toward Sustainability-Aware LLM Inference on Edge Clusters
Kolichala Rajashekar, Nafiseh Sharghivand, Radu Prodan +1
Large language models (LLMs) require substantial computational resources, leading to significant carbon emissions and operational costs. Although training is energy-intensive, the…