activity
20172025
most citedSound Event Detection with Binary Neural Networks on Tightly Power-Constrained IoT Devices

39 citations · 84 across the 23 of their papers we have counts for

collaborators
Showing cs.ARShow all

6 papers · 1 filter

cs.AR2023

Stella Nera: A Differentiable Maddness-Based Hardware Accelerator for Efficient Approximate Matrix Multiplication

Jannis Schönleber, Lukas Cavigelli, Matteo Perotti +2

Artificial intelligence has surged in recent years, with advancements in machine learning rapidly impacting nearly every area of life. However, the growing complexity of these mode…

cs.AR2023

Ara2: Exploring Single- and Multi-Core Vector Processing with an Efficient RVV 1.0 Compliant Open-Source Processor

Matteo Perotti, Matheus Cavalcante, Renzo Andri +2

Vector processing is highly effective in boosting processor performance and efficiency for data-parallel workloads. In this paper, we present Ara2, the first fully open-source vect…

cs.AR2023

ReDSEa: Automated Acceleration of Triangular Solver on Supercloud Heterogeneous Systems

Georgios Zacharopoulos, Ilias Bournias, Verner Vlacic +1

When utilized effectively, Supercloud heterogeneous systems have the potential to significantly enhance performance. Our ReDSEa tool-chain automates the mapping, load balancing, sc…

cs.AR20232 cited

Flex-SFU: Accelerating DNN Activation Functions by Non-Uniform Piecewise Approximation

Enrico Reggiani, Renzo Andri, Lukas Cavigelli

Modern DNN workloads increasingly rely on activation functions consisting of computationally complex operations. This poses a challenge to current accelerators optimized for convol…

cs.AR2022

Going Further With Winograd Convolutions: Tap-Wise Quantization for Efficient Inference on 4x4 Tile

Renzo Andri, Beatrice Bussolino, Antonio Cipolletta +2

Most of today's computer vision pipelines are built around deep neural networks, where convolution operations require most of the generally high compute effort. The Winograd convol…

cs.AR2020

CUTIE: Beyond PetaOp/s/W Ternary DNN Inference Acceleration with Better-than-Binary Energy Efficiency

Moritz Scherer, Georg Rutishauser, Lukas Cavigelli +1

We present a 3.1 POp/s/W fully digital hardware accelerator for ternary neural networks. CUTIE, the Completely Unrolled Ternary Inference Engine, focuses on minimizing non-computat…