9 citations · 14 across the 9 of their papers we have counts for
6 papers · 1 filter
Exploiting pre-optimized kernels with polyhedral transformations for CGRA compilation
Yuxuan Wang, María José Belda, Fernando Castro +3
Modern computing workloads commonly involve matrix-matrix multiplication (mmul) as a core computing pattern. Coarse-Grained Reconfigurable Arrays (CGRAs) can flexibly and efficient…
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design
Qunyou Liu, Marina Zapater, David Atienza
Transformers have revolutionized AI in natural language processing and computer vision, but their large computation and memory demands pose major challenges for hardware accelerati…
A flexible framework for early power and timing comparison of time-multiplexed CGRA kernel executions
Maxime Henri Aspros, Juan Sapriza, Giovanni Ansaloni +1
At the intersection between traditional CPU architectures and more specialized options such as FPGAs or ASICs lies the family of reconfigurable hardware architectures, termed Coars…
MatrixFlow: System-Accelerator co-design for high-performance transformer applications
Qunyou Liu, Marina Zapater, David Atienza
Transformers are central to advances in artificial intelligence (AI), excelling in fields ranging from computer vision to natural language processing. Despite their success, their…
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
Qunyou Liu, Marina Zapater, David Atienza
The growing demand for efficient, high-performance processing in machine learning (ML) and image processing has made hardware accelerators, such as GPUs and Data Streaming Accelera…
Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge Systems
Pedro Palacios, Rafael Medina, Jean-Luc Rouas +2
Efficient deployment of resource-intensive transformers on edge devices necessitates cross-stack optimization. We thus study the interrelation between structured pruning and systol…