3 papers
cs.DC2026
[AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent Workflows
Arup Kumar Sarker, Alexander James Halpern, Mills Staylor +5
The paper presents AAFLOW+, a framework that treats key‑value (KV) caches as distributed objects, enabling zero‑copy sharing of model state across multi‑agent LLM workflows to cut…
cs.DC2026
Combining Serverless and High-Performance Computing Paradigms to support ML Data-Intensive Applications
Mills Staylor, Arup Kumar Sarker, Gregor von Laszewski +3
Data is found everywhere, from health and human infrastructure to the surge of sensors and the proliferation of internet-connected devices. To meet this challenge, the data enginee…
cs.LG2024
Ensuring Fair LLM Serving Amid Diverse Applications
Redwan Ibne Seraj Khan, Kunal Jain, Haiying Shen +12
In a multi-tenant large language model (LLM) serving platform hosting diverse applications, some users may submit an excessive number of requests, causing the service to become una…