collaborators

5 papers

cs.AR2025

Squire: A General-Purpose Accelerator to Exploit Fine-Grain Parallelism on Dependency-Bound Kernels

Rubén Langarita, Jesús Alastruey-Benedé, Pablo Ibáñez-Marín +3

Multiple HPC applications are often bottlenecked by compute-intensive kernels implementing complex dependency patterns (data-dependency bound). Traditional general-purpose accelera…

cs.AR2025

A Flexible Instruction Set Architecture for Efficient GEMMs

Alexandre de Limas Santana, Adrià Armejach, Francesc Martinez +2

GEneral Matrix Multiplications (GEMMs) are recurrent in high-performance computing and deep learning workloads. Typically, high-end CPUs accelerate GEMM workloads with Single-Instr…

cs.AR2025

Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures

Marco Siracusa, Olivia Hsu, Victor Soria-Pardos +8

Irregular embedding lookups are a critical bottleneck in recommender models, sparse large language models, and graph learning models. In this paper, we first demonstrate that, by o…

cs.DB2025

On Usage of Non-Volatile Memory as Primary Storage for Database Management Systems

Naveed Ul Mustafa, Adri`a Armejach, Ozcan Ozturk +2

This paper explores the implications of employing non-volatile memory (NVM) as primary storage for a data base management system (DBMS). We investigate the modifications necessary…

cs.AR2024

A Mess of Memory System Benchmarking, Simulation and Application Profiling

Pouya Esmaili-Dokht, Francesco Sgherzi, Valeria Soldera Girelli +15

The Memory stress (Mess) framework provides a unified view of the memory system benchmarking, simulation and application profiling. The Mess benchmark provides a holistic and detai…