4 papers
SMaRTT: Sender-based Marked Rapidly-adapting Trimmed & Timed Transport
Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini +10
With the rapid growth of artificial intelligence (AI) workloads in datacenters, the Ultra Ethernet Consortium (UEC) has defined a new high-performance transport layer to deliver th…
MLIR-Forge: A Modular Framework for Language Smiths
Berke Ates, Philipp Schaad, Timo Schneider +2
Optimizing compilers are essential for the efficient and correct execution of software across various scientific fields. Domain-specific languages (DSL) typically use higher level…
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
Siyuan Shen, Mikhail Khalilov, Lukas Gianinazzi +6
Resource disaggregation is a promising technique for improving the efficiency of large-scale computing systems. However, this comes at the cost of increased memory access latency d…
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
Andrei Ivanov, Siyuan Shen, Gioele Gottardo +5
The increasing complexity of machine learning models and the proliferation of diverse hardware architectures (CPUs, GPUs, accelerators) make achieving optimal performance a signifi…