4 papers
A flexible framework for early power and timing comparison of time-multiplexed CGRA kernel executions
Maxime Henri Aspros, Juan Sapriza, Giovanni Ansaloni +1
At the intersection between traditional CPU architectures and more specialized options such as FPGAs or ASICs lies the family of reconfigurable hardware architectures, termed Coars…
MatrixFlow: System-Accelerator co-design for high-performance transformer applications
Qunyou Liu, Marina Zapater, David Atienza
Transformers are central to advances in artificial intelligence (AI), excelling in fields ranging from computer vision to natural language processing. Despite their success, their…
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
Qunyou Liu, Marina Zapater, David Atienza
The growing demand for efficient, high-performance processing in machine learning (ML) and image processing has made hardware accelerators, such as GPUs and Data Streaming Accelera…
Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge Systems
Pedro Palacios, Rafael Medina, Jean-Luc Rouas +2
Efficient deployment of resource-intensive transformers on edge devices necessitates cross-stack optimization. We thus study the interrelation between structured pruning and systol…