8 papers
An MLIR-Based Compilation Framework for Control Flow Management on Coarse Grained Reconfigurable Arrays
Yuxuan Wang, Cristian Tirelli, Giovanni Ansaloni +2
Coarse Grained Reconfigurable Arrays (CGRAs) present both high flexibility and efficiency, making them well-suited for the acceleration of intensive workloads. Nevertheless, a key…
Exploiting pre-optimized kernels with polyhedral transformations for CGRA compilation
Yuxuan Wang, MarÃa José Belda, Fernando Castro +3
Modern computing workloads commonly involve matrix-matrix multiplication (mmul) as a core computing pattern. Coarse-Grained Reconfigurable Arrays (CGRAs) can flexibly and efficient…
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design
Qunyou Liu, Marina Zapater, David Atienza
Transformers have revolutionized AI in natural language processing and computer vision, but their large computation and memory demands pose major challenges for hardware accelerati…
A flexible framework for early power and timing comparison of time-multiplexed CGRA kernel executions
Maxime Henri Aspros, Juan Sapriza, Giovanni Ansaloni +1
At the intersection between traditional CPU architectures and more specialized options such as FPGAs or ASICs lies the family of reconfigurable hardware architectures, termed Coars…
Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge Systems
Pedro Palacios, Rafael Medina, Jean-Luc Rouas +2
Efficient deployment of resource-intensive transformers on edge devices necessitates cross-stack optimization. We thus study the interrelation between structured pruning and systol…
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
Qunyou Liu, Marina Zapater, David Atienza
The growing demand for efficient, high-performance processing in machine learning (ML) and image processing has made hardware accelerators, such as GPUs and Data Streaming Accelera…