3 papers
cs.AR2026
Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era
Junrui Pan, Weili An, Cesar Avalos Baddouh +9
The rapid evolution of machine learning workloads has fundamentally transformed GPU hardware, driving architectures toward Multi-Chip Module (MCM) topologies, asynchronous executio…
cs.PL2026
Triton for MTIA: Bridging the Programming Model Gaps for Custom AI Accelerators
Haishan Zhu, Domi Yan, Michael Levesque-Dion +40
The rapid growth in machine learning workloads has fueled the proliferation of custom accelerator architectures. Designed from the ground up, these accelerators often expose progra…
cs.AR2024
RayFlex: An Open-Source RTL Implementation of the Hardware Ray Tracer Datapath
Fangjia Shen, Aaron Barnes, Anusuya Nallathambi +1
The advent of hardware ray tracing (RT) units has brought unprecedented realism to real-time rendered computer graphics. However, the potential of these units extends beyond graphi…