collaborators

11 papers

cs.DC2026

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving

Shrikara Arun, Anjaly Parayil, Srikant Bharadwaj +2

Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates a new asymmetry: under bursty, heavy-tail…

cs.LG2026

Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems

Fiza Husain, Anson Bastos, Anjaly Parayil +4

In this paper, we present DiRecGNN, an attention-enhanced entity recommendation framework for monitoring cloud services at Microsoft. We provide insights on the usefulness of this…

cs.DC2026

CWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms

Tella Rajashekhar Reddy, Atharva Deshmukh, Liangcheng Yu +7

AI power demand is growing at an unprecedented rate while power grids are often ailing and struggle to keep up. Grid expansion comes with high capital expenditure and long-distance…

cs.DC2026

Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference

Anish Biswas, Kanishk Goel, Srivarshinee S +5

Agentic applications are LLMs that iteratively invoke external tools to accomplish complex tasks. Such tool-based agents are rapidly becoming the dominant paradigm for deploying la…

cs.DC2026

A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring

Anson Bastos, Shreeya Venneti, Anjaly Parayil +3

Reliability of large-scale cloud services is critical for user satisfaction and business continuity. Despite significant investments in reliability engineering, production incident…

cs.DC2025

Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems

Shashwat Jaiswal, Shrikara Arun, Anjaly Parayil +8

Low-Rank Adaptation (LoRA) has become the de facto method for parameter-efficient fine-tuning of large language models (LLMs), enabling rapid adaptation to diverse domains. In prod…