1 citations · 1 across the 3 of their papers we have counts for
3 papers
Ascendra: Dynamic Request Prioritization for Efficient LLM Serving
Azam Ikram, Xiang Li, Sameh Elnikety +1
The rapid advancement of Large Language Models (LLMs) has driven the need for more efficient serving strategies. In this context, efficiency refers to the proportion of requests th…
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
Enrique Saurez, Joshua Fried, Gohar Irfan Chaudhry +5
This report explores the use of kernel-bypass networking in FaaS runtimes and demonstrates how using Junction, a novel kernel-bypass system, as the backend for executing components…
Analytically-Driven Resource Management for Cloud-Native Microservices
Yanqi Zhang, Zhuangzhuang Zhou, Sameh Elnikety +1
Resource management for cloud-native microservices has attracted a lot of recent attention. Previous work has shown that machine learning (ML)-driven approaches outperform traditio…