2 papers
cs.AR2026
Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era
Junrui Pan, Weili An, Cesar Avalos Baddouh +9
The rapid evolution of machine learning workloads has fundamentally transformed GPU hardware, driving architectures toward Multi-Chip Module (MCM) topologies, asynchronous executio…
cs.AR2023
LRMP: Layer Replication with Mixed Precision for Spatial In-memory DNN Accelerators
Abinand Nallathambi, Christin David Bose, Wilfried Haensch +1
In-memory computing (IMC) with non-volatile memories (NVMs) has emerged as a promising approach to address the rapidly growing computational demands of Deep Neural Networks (DNNs).…