collaborators

8 papers

cs.AR2025

TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads

Navaneeth Kunhi Purayil, Diyou Shen, Matteo Perotti +1

The fast evolution of Machine Learning (ML) models requires flexible and efficient hardware solutions as hardwired accelerators face rapid obsolescence. Vector processors are fully…

cs.AR2025

Stella Nera: A Differentiable Maddness-Based Hardware Accelerator for Efficient Approximate Matrix Multiplication

Jannis Schönleber, Lukas Cavigelli, Matteo Perotti +2

Artificial intelligence has surged in recent years, with advancements in machine learning rapidly impacting nearly every area of life. However, the growing complexity of these mode…

cs.AR2025

AraOS: Analyzing the Impact of Virtual Memory Management on Vector Unit Performance

Matteo Perotti, Vincenzo Maisto, Moritz Imfeld +3

Vector processor architectures offer an efficient solution for accelerating data-parallel workloads (e.g., ML, AI), reducing instruction count, and enhancing processing efficiency.…

cs.AR2025

Quadrilatero: A RISC-V programmable matrix coprocessor for low-power edge applications

Danilo Cammarata, Matteo Perotti, Marco Bertuletti +4

The rapid growth of AI-based Internet-of-Things applications increased the demand for high-performance edge processing engines on a low-power budget and tight area constraints. As…

cs.AR2025

A Reliable, Time-Predictable Heterogeneous SoC for AI-Enhanced Mixed-Criticality Edge Applications

Angelo Garofalo, Alessandro Ottaviano, Matteo Perotti +20

Next-generation mixed-criticality Systems-on-chip (SoCs) for robotics, automotive, and space must execute mixed-criticality AI-enhanced sensor processing and control workloads, ens…

cs.AR2025

AraXL: A Physically Scalable, Ultra-Wide RISC-V Vector Processor Design for Fast and Efficient Computation on Long Vectors

Navaneeth Kunhi Purayil, Matteo Perotti, Tim Fischer +1

The ever-growing scale of data parallelism in today's HPC and ML applications presents a big challenge for computing architectures' energy efficiency and performance. Vector proces…