11 citations
- Microsoft Research (United Kingdom)GB4 papers
- Carnegie Mellon UniversityUS2 papers
- Microsoft Research Asia (China)CN2 papers
- Microsoft Research (India)IN2 papers
- Microsoft Research New York City (United States)2 papers
- Princeton UniversityUS2 papers
- The University of Texas at AustinUS2 papers
- Tsinghua UniversityCN2 papers
- University of ChicagoUS2 papers
- William & MaryUS2 papers
- Allergan (India)IN1 paper
- Cisco Systems (United States)US1 paper
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026★ 3 cited
Energy Use of AI Inference, Efficiency Pathways, and Test-Time Scaling
Felipe Oviedo, Fiodar Kazhamiaka, Esha Choukse +5
As AI inference scales to billions of queries, estimates of per-query energy use are increasingly important for capacity planning, efficiency interventions, and policy. Yet many pu…
cs.LG2026★ 1 cited
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
Yaoqi Chen, Jinkai Zhang, Baotong Lu +16
Recent large language models (LLMs) are rapidly extending their context windows, yet inference throughput lags due to increasing GPU memory and bandwidth demands. This is because t…