5 papers
CXLRAMSim v1.0: System-Level Exploration of CXL Memory Expander Cards
Karan Pathak, David Atienza, Marina Zapater
The growing demands in the training and inference of Large Language Models (LLMs) are accelerating the adoption of scale-up systems that extend server shared memory through the use…
CloudFormer: An Attention-based Performance Prediction for Public Clouds with Unknown Workload
Amirhossein Shahbazinia, Darong Huang, Luis Costero +1
Cloud platforms are increasingly relied upon to host diverse, resource-intensive workloads due to their scalability, flexibility, and cost-efficiency. In multi-tenant cloud environ…
3D-ICE 4.0: Accurate and efficient thermal modeling for 2.5D/3D heterogeneous chiplet systems
Kai Zhu, Darong Huang, Luis Costero +1
The increasing power densities and intricate heat dissipation paths in advanced 2.5D/3D chiplet systems necessitate thermal modeling frameworks that deliver detailed thermal maps w…
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
Qunyou Liu, Darong Huang, Marina Zapater +1
Large Language Models (LLMs) are becoming the backbone of modern cloud services, yet their inference costs are dominated by GPU energy. Unlike traditional GPU workloads, LLM infere…
A 20-Year Retrospective on Power and Thermal Modeling and Management
David Atienza, Kai Zhu, Darong Huang +1
As processor performance advances, increasing power densities and complex thermal behaviors threaten both energy efficiency and system reliability. This survey covers more than two…