14 citations · 19 across the 3 of their papers we have counts for
3 papers
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
Jovan Stojkovic, Esha Choukse, Chaojie Zhang +2
With the ubiquitous use of modern large language models (LLMs) across industries, the inference serving for these models is ever expanding. Given the high compute and memory requir…
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
Enrique Saurez, Joshua Fried, Gohar Irfan Chaudhry +5
This report explores the use of kernel-bypass networking in FaaS runtimes and demonstrates how using Junction, a novel kernel-bypass system, as the backend for executing components…
POLCA: Power Oversubscription in LLM Cloud Providers
Pratyush Patel, Esha Choukse, Chaojie Zhang +4
Recent innovation in large language models (LLMs), and their myriad use-cases have rapidly driven up the compute capacity demand for datacenter GPUs. Several cloud providers and ot…