5 papers
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
Jonas Svedas, Nathan Laubeuf, Ryan Harvey +6
Predicting the performance of large-scale distributed machine learning (ML) workloads across multiple accelerator architectures remains a central challenge in ML system design. Exi…
Decoupled Control Flow and Data Access in RISC-V GPGPUs
Giuseppe M. Sarda, Nimish Shah, Abubakr Nada +2
Vortex, a newly proposed open-source GPGPU platform based on the RISC-V ISA, offers a valid alternative for GPGPU research over the broadly-used modeling platforms based on commerc…
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
Jonas Svedas, Hannah Watson, Nathan Laubeuf +6
Distributed deep neural networks (DNNs) have become a cornerstone for scaling machine learning to meet the demands of increasingly complex applications. However, the rapid growth i…
Addressing memory bandwidth scalability in vector processors for streaming applications
Jordi Altayo, Paul Delestrac, David Novo +3
As the size of artificial intelligence and machine learning (AI/ML) models and datasets grows, the memory bandwidth becomes a critical bottleneck. The paper presents a novel extend…
A System Level Performance Evaluation for Superconducting Digital Systems
Joyjit Kundu, Debjyoti Bhattacharjee, Nathan Josephsen +8
Superconducting Digital (SCD) technology offers significant potential for enhancing the performance of next generation large scale compute workloads. By leveraging advanced lithogr…