most citedOptimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

cs.DC2025

A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO

Jonas Svedas, Hannah Watson, Nathan Laubeuf +6

Distributed deep neural networks (DNNs) have become a cornerstone for scaling machine learning to meet the demands of increasingly complex applications. However, the rapid growth i…

cs.AR2025

Addressing memory bandwidth scalability in vector processors for streaming applications

Jordi Altayo, Paul Delestrac, David Novo +3

As the size of artificial intelligence and machine learning (AI/ML) models and datasets grows, the memory bandwidth becomes a critical bottleneck. The paper presents a novel extend…

cs.LO2025

Linear Decomposition of the Majority Boolean Function using the Ones on Smaller Variables

Anupam Chattopadhyay, Debjyoti Bhattacharjee, Subhamoy Maitra

A long-investigated problem in circuit complexity theory is to decompose an -input or -variable Majority Boolean function (call it ) using -input ones (), $k < n…

cs.LG2024

SAfEPaTh: A System-Level Approach for Efficient Power and Thermal Estimation of Convolutional Neural Network Accelerator

Yukai Chen, Simei Yang, Debjyoti Bhattacharjee +2

The design of energy-efficient, high-performance, and reliable Convolutional Neural Network (CNN) accelerators involves significant challenges due to complex power and thermal mana…

cs.AR20241 cited

Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis

Giuseppe M. Sarda, Nimish Shah, Debjyoti Bhattacharjee +2

GPGPU execution analysis has always been tied to closed-source, proprietary benchmarking tools that provide high-level, non-exhaustive, and/or statistical information, preventing a…