most citedTowards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference

14 citations · 21 across the 5 of their papers we have counts for

collaborators

5 papers

cs.DC20242 cited

Workload Intelligence: Punching Holes Through the Cloud Abstraction

Lexiang Huang, Anjaly Parayil, Jue Zhang +13

Today, cloud workloads are essentially opaque to the cloud platform. Typically, the only information the platform receives is the virtual machine (VM) type and possibly a decoratio…

cs.AI202414 cited

Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference

Jovan Stojkovic, Esha Choukse, Chaojie Zhang +2

With the ubiquitous use of modern large language models (LLMs) across industries, the inference serving for these models is ever expanding. Given the high compute and memory requir…

cs.DC2024

Junctiond: Extending FaaS Runtimes with Kernel-Bypass

Enrique Saurez, Joshua Fried, Gohar Irfan Chaudhry +5

This report explores the use of kernel-bypass networking in FaaS runtimes and demonstrates how using Junction, a novel kernel-bypass system, as the backend for executing components…

cs.HC2024

Risk-aware Adaptive Virtual CPU Oversubscription in Microsoft Cloud via Prototypical Human-in-the-loop Imitation Learning

Lu Wang, Mayukh Das, Fangkai Yang +11

Oversubscription is a prevalent practice in cloud services where the system offers more virtual resources, such as virtual cores in virtual machines, to users or applications than…

cs.DC20235 cited

POLCA: Power Oversubscription in LLM Cloud Providers

Pratyush Patel, Esha Choukse, Chaojie Zhang +4

Recent innovation in large language models (LLMs), and their myriad use-cases have rapidly driven up the compute capacity demand for datacenter GPUs. Several cloud providers and ot…