collaborators

5 papers

cs.LG2026

Energy Use of AI Inference, Efficiency Pathways, and Test-Time Scaling

Felipe Oviedo, Fiodar Kazhamiaka, Esha Choukse +5

As AI inference scales to billions of queries, estimates of per-query energy use are increasingly important for capacity planning, efficiency interventions, and policy. Yet many pu…

cs.AI2026

Natural Language Query to Configuration for Retrieval Agents

Melissa Z. Pan, Negar Arabzadeh, Mathew Jacob +3

Modern retrieval agents expose many configuration choices -- LLM, retriever, number of documents, number of hops, and synthesis strategy -- each shaping both answer quality and ser…

cs.DC2026

Designing Datacenter Power Delivery Hierarchies for the AI Era

Grant Wilkins, Fiodar Kazhamiaka, Alok Gautam Kumbhare +2

Demand for AI accelerators is rapidly increasing rack power density, with projections approaching 1MW per deployment by 2027. This poses a major challenge for datacenter power deli…

cs.AR2026

Octopus: Enhancing CXL Memory Pods via Sparse Topology

Yuhong Zhong, Fiodar Kazhamiaka, Pantea Zardoshti +4

The Compute Express Link (CXL) interconnect enables compute "pods" that pool memory across servers to reduce cost and improve efficiency. These pods also facilitate pairwise commun…

cs.DC2026

From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning

Grant Wilkins, Fiodar Kazhamiaka, Ram Rajagopal

Datacenter operators and electrical utilities rely on power traces at different spatiotemporal scales. Operators use fine-grained traces for provisioning, facility management, and…