5 papers
Squire: A General-Purpose Accelerator to Exploit Fine-Grain Parallelism on Dependency-Bound Kernels
Rubén Langarita, Jesús Alastruey-Benedé, Pablo Ibáñez-MarÃn +3
Multiple HPC applications are often bottlenecked by compute-intensive kernels implementing complex dependency patterns (data-dependency bound). Traditional general-purpose accelera…
A Flexible Instruction Set Architecture for Efficient GEMMs
Alexandre de Limas Santana, Adrià Armejach, Francesc Martinez +2
GEneral Matrix Multiplications (GEMMs) are recurrent in high-performance computing and deep learning workloads. Typically, high-end CPUs accelerate GEMM workloads with Single-Instr…
Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures
Marco Siracusa, Olivia Hsu, Victor Soria-Pardos +8
Irregular embedding lookups are a critical bottleneck in recommender models, sparse large language models, and graph learning models. In this paper, we first demonstrate that, by o…
On Usage of Non-Volatile Memory as Primary Storage for Database Management Systems
Naveed Ul Mustafa, Adri`a Armejach, Ozcan Ozturk +2
This paper explores the implications of employing non-volatile memory (NVM) as primary storage for a data base management system (DBMS). We investigate the modifications necessary…
A Mess of Memory System Benchmarking, Simulation and Application Profiling
Pouya Esmaili-Dokht, Francesco Sgherzi, Valeria Soldera Girelli +15
The Memory stress (Mess) framework provides a unified view of the memory system benchmarking, simulation and application profiling. The Mess benchmark provides a holistic and detai…