most citedDARKSIDE: A Heterogeneous RISC-V Compute Cluster for Extreme-Edge On-Chip DNN Inference and Training

34 citations · 57 across the 3 of their papers we have counts for

collaborators
Showing cs.ARShow all

5 papers · 1 filter

cs.AR2024

MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication

Matteo Perotti, Yichao Zhang, Matheus Cavalcante +2

Dense Matrix Multiplication (MatMul) is arguably one of the most ubiquitous compute-intensive kernels, spanning linear algebra, DSP, graphics, and machine learning applications. Th…

cs.AR2023

Near-Memory Parallel Indexing and Coalescing: Enabling Highly Efficient Indirect Access for SpMV

Chi Zhang, Paul Scheffler, Thomas Benz +2

Sparse matrix vector multiplication (SpMV) is central to numerous data-intensive applications, but requires streaming indirect memory accesses that severely degrade both processing…

cs.AR202334 cited

DARKSIDE: A Heterogeneous RISC-V Compute Cluster for Extreme-Edge On-Chip DNN Inference and Training

Angelo Garofalo, Yvan Tortorella, Matteo Perotti +5

On-chip DNN inference and training at the Extreme-Edge (TinyML) impose strict latency, throughput, accuracy and flexibility requirements. Heterogeneous clusters are promising solut…

cs.AR2023

Quark: An Integer RISC-V Vector Processor for Sub-Byte Quantized DNN Inference

MohammadHossein AskariHemmat, Theo Dupuis, Yoan Fournier +8

In this paper, we present Quark, an integer RISC-V vector processor specifically tailored for sub-byte DNN inference. Quark is implemented in GlobalFoundries' 22FDX FD-SOI technolo…

cs.AR202223 cited

Spatz: A Compact Vector Processing Unit for High-Performance and Energy-Efficient Shared-L1 Clusters

Matheus Cavalcante, Domenic Wüthrich, Matteo Perotti +2

While parallel architectures based on clusters of Processing Elements (PEs) sharing L1 memory are widespread, there is no consensus on how lean their PE should be. Architecting PEs…