2 papers
cs.AR2026
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
Danilo Cammarata, Angelo Garofalo, Luca Benini
General matrix multiply (GEMM) dominates both execution time and energy consumption of modern machine learning (ML) workloads, placing increasing pressure on hardware efficiency. W…
cs.AR2025
Quadrilatero: A RISC-V programmable matrix coprocessor for low-power edge applications
Danilo Cammarata, Matteo Perotti, Marco Bertuletti +4
The rapid growth of AI-based Internet-of-Things applications increased the demand for high-performance edge processing engines on a low-power budget and tight area constraints. As…