1 citations · 1 across the 3 of their papers we have counts for
4 papers
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
JooYoung Park, Kevin Nguetchouang, Jovan Stojkovic +4
Serverless computing relies on extreme multi-tenancy to remain economically viable, driving providers to rely on virtual machines (VMs) that ensure strong isolation and seamless ec…
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri +1
The rapid rise of large language models (LLMs) has been driving an enormous demand for AI inference infrastructure, mainly powered by high-end GPUs. While these accelerators offer…
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri +5
The rising demand for generative large language models (LLMs) poses challenges for thermal and power management in cloud datacenters. Traditional techniques often are inadequate fo…
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
Nikoleta Iliakopoulou, Jovan Stojkovic, Chloe Alverti +3
The widespread adoption of LLMs has driven an exponential rise in their deployment, imposing substantial demands on inference clusters. These clusters must handle numerous concurre…