activity
20242026
collaborators

5 papers

cs.DC2026

Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO

Jonas Svedas, Nathan Laubeuf, Ryan Harvey +6

Predicting the performance of large-scale distributed machine learning (ML) workloads across multiple accelerator architectures remains a central challenge in ML system design. Exi…

cs.AR2025

Decoupled Control Flow and Data Access in RISC-V GPGPUs

Giuseppe M. Sarda, Nimish Shah, Abubakr Nada +2

Vortex, a newly proposed open-source GPGPU platform based on the RISC-V ISA, offers a valid alternative for GPGPU research over the broadly-used modeling platforms based on commerc…

cs.DC2025

A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO

Jonas Svedas, Hannah Watson, Nathan Laubeuf +6

Distributed deep neural networks (DNNs) have become a cornerstone for scaling machine learning to meet the demands of increasingly complex applications. However, the rapid growth i…

cs.AR2025

Addressing memory bandwidth scalability in vector processors for streaming applications

Jordi Altayo, Paul Delestrac, David Novo +3

As the size of artificial intelligence and machine learning (AI/ML) models and datasets grows, the memory bandwidth becomes a critical bottleneck. The paper presents a novel extend…

cs.AR2024

A System Level Performance Evaluation for Superconducting Digital Systems

Joyjit Kundu, Debjyoti Bhattacharjee, Nathan Josephsen +8

Superconducting Digital (SCD) technology offers significant potential for enhancing the performance of next generation large scale compute workloads. By leveraging advanced lithogr…