4 papers · 1 filter
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
Endri Taka, Dimitrios Gourounas, Andreas Gerstlauer +2
FPGAs are a promising platform for accelerating Deep Learning (DL) applications, due to their high performance, low power consumption, and reconfigurability. Recently, the leading…
Lightweight ML-based Runtime Prefetcher Selection on Many-core Platforms
Erika S. Alcorta, Mahesh Madhav, Scott Tetrick +2
Modern computer designs support composite prefetching, where multiple individual prefetcher components are used to target different memory access patterns. However, multiple prefet…
Virtual-Link: A Scalable Multi-Producer, Multi-Consumer Message Queue Architecture for Cross-Core Communication
Qinzhe Wu, Jonathan Beard, Ashen Ekanayake +2
Cross-core communication is increasingly a bottleneck as the number of processing elements increase per system-on-chip. Typical hardware solutions to cross-core communication are o…
Exploiting Errors for Efficiency: A Survey from Circuits to Algorithms
Phillip Stanley-Marbell, Armin Alaghi, Michael Carbin +13
When a computational task tolerates a relaxation of its specification or when an algorithm tolerates the effects of noise in its execution, hardware, programming languages, and sys…