2 papers
cs.ET2026
Beyond HBM-on-GPU: Thermal Design Envelope for 3D Volumetric DRAM-on-GPU Integration
Yukai Chen, Melina Lofrano, Khakim Akhunov +12
The scaling of GPUs for AI and HPC workloads is increasingly constrained by the capacity, bandwidth, and thermal limits of both 2.5D HBM-GPU and direct-stacked 3D HBM-on-GPU integr…
cs.DC2026
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
Jonas Svedas, Nathan Laubeuf, Ryan Harvey +6
Predicting the performance of large-scale distributed machine learning (ML) workloads across multiple accelerator architectures remains a central challenge in ML system design. Exi…