4 papers · 1 filter
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
JooYoung Park, Kevin Nguetchouang, Jovan Stojkovic +4
Serverless computing relies on extreme multi-tenancy to remain economically viable, driving providers to rely on virtual machines (VMs) that ensure strong isolation and seamless ec…
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
Nikoleta Iliakopoulou, Jovan Stojkovic, Chloe Alverti +3
The widespread adoption of LLMs has driven an exponential rise in their deployment, imposing substantial demands on inference clusters. These clusters must handle numerous concurre…
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
Jovan Stojkovic, Chaojie Zhang, Ãñigo Goiri +5
The rising demand for generative large language models (LLMs) poses challenges for thermal and power management in cloud datacenters. Traditional techniques often are inadequate fo…
Workload Intelligence: Punching Holes Through the Cloud Abstraction
Lexiang Huang, Anjaly Parayil, Jue Zhang +13
Today, cloud workloads are essentially opaque to the cloud platform. Typically, the only information the platform receives is the virtual machine (VM) type and possibly a decoratio…