activity
20162025
most citedLightweight Neural Architecture Search for Temporal Convolutional Networks at the Edge

48 citations · 92 across the 13 of their papers we have counts for

collaborators
Showing cs.ARShow all

7 papers · 1 filter

cs.AR2025

Work-In-Progress: Accelerating Numpy With OpenBLAS For Open-Source RISC-V Chips

Cyril Koenig, Enrico Zelioli, Frank K. Gürkaynak +1

RISC-V allows for building general-purpose computing platforms with programmable accelerators around a single open-source ISA. However, leveraging heterogeneous SoCs within high-le…

cs.AR20242 cited

Insights from Basilisk: Are Open-Source EDA Tools Ready for a Multi-Million-Gate, Linux-Booting RV64 SoC Design?

Philippe Sauter, Thomas Benz, Paul Scheffler +2

Designing complex, multi-million-gate application-specific integrated circuits requires robust and mature electronic design automation (EDA) tools. We describe our efforts in enhan…

cs.AR20231 cited

ColibriES: A Milliwatts RISC-V Based Embedded System Leveraging Neuromorphic and Neural Networks Hardware Accelerators for Low-Latency Closed-loop Control Applications

Georg Rutishauser, Robin Hunziker, Alfio Di Mauro +3

End-to-end event-based computation has the potential to push the envelope in latency and energy efficiency for edge AI applications. Unfortunately, event-based sensors (e.g., DVS c…

cs.AR2023

Quark: An Integer RISC-V Vector Processor for Sub-Byte Quantized DNN Inference

MohammadHossein AskariHemmat, Theo Dupuis, Yoan Fournier +8

In this paper, we present Quark, an integer RISC-V vector processor specifically tailored for sub-byte DNN inference. Quark is implemented in GlobalFoundries' 22FDX FD-SOI technolo…

cs.AR2022

Soft Tiles: Capturing Physical Implementation Flexibility for Tightly-Coupled Parallel Processing Clusters

Gianna Paulin, Matheus Cavalcante, Paul Scheffler +4

Modern high-performance computing architectures (Multicore, GPU, Manycore) are based on tightly-coupled clusters of processing elements, physically implemented as rectangular tiles…

cs.AR202223 cited

Spatz: A Compact Vector Processing Unit for High-Performance and Energy-Efficient Shared-L1 Clusters

Matheus Cavalcante, Domenic Wüthrich, Matteo Perotti +2

While parallel architectures based on clusters of Processing Elements (PEs) sharing L1 memory are widespread, there is no consensus on how lean their PE should be. Architecting PEs…