most citedTAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms

1 citations · 2 across the 4 of their papers we have counts for

collaborators

8 papers

cs.MA2025

Sherlock: Reliable and Efficient Agentic Workflow Execution

Yeonju Ro, Haoran Qiu, Íñigo Goiri +6

With the increasing adoption of large language models (LLM), agentic workflows, which compose multiple LLM calls with tools, retrieval, and reasoning steps, are increasingly replac…

cs.AI2025

Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework

Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri +1

The rapid rise of large language models (LLMs) has been driving an enormous demand for AI inference infrastructure, mainly powered by high-end GPUs. While these accelerators offer…

cs.MA2025

Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms

Gohar Irfan Chaudhry, Esha Choukse, Haoran Qiu +4

Agentic workflows commonly coordinate multiple models and tools with complex control logic. They are quickly becoming the dominant paradigm for AI applications. However, serving th…

cs.AR20251 cited

Power Stabilization for AI Training Datacenters

Esha Choukse, Brijesh Warrier, Scot Heath +54

Large Artificial Intelligence (AI) training workloads spanning several tens of thousands of GPUs present unique power management challenges. These arise due to the high variability…

cs.DC20251 cited

TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms

Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri +5

The rising demand for generative large language models (LLMs) poses challenges for thermal and power management in cloud datacenters. Traditional techniques often are inadequate fo…

cs.DC2025

Towards Resource-Efficient Compound AI Systems

Gohar Irfan Chaudhry, Esha Choukse, Íñigo Goiri +3

Compound AI Systems, integrating multiple interacting components like models, retrievers, and external tools, have emerged as essential for addressing complex AI tasks. However, cu…