3 papers
cs.DC2026
Sangam: Efficiently Serving Diffusion LLMs with the AR Stack
Nitin Kedia, Saurabh Agarwal, Myungjin Lee +1
Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation. Their bidirectional attent…
cs.DC2026
Nalar: An agent serving framework
Marco Laju, Donghyun Son, Saurabh Agarwal +4
LLM-driven agentic applications increasingly automate complex, multi-step tasks, but serving them efficiently remains challenging due to heterogeneous components, dynamic and model…
cs.LG2025
On Evaluating Performance of LLM Inference Serving Systems
Amey Agrawal, Nitin Kedia, Anmol Agarwal +5
The rapid evolution of Large Language Model (LLM) inference systems has yielded significant efficiency improvements. However, our systematic analysis reveals that current evaluatio…