2 papers
cs.AR2026
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
Danilo Cammarata, Angelo Garofalo, Luca Benini
General matrix multiply (GEMM) dominates both execution time and energy consumption of modern machine learning (ML) workloads, placing increasing pressure on hardware efficiency. W…
cs.DC2026
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
Enrico Russo, Mohamed Amine Hamdi, Alessandro Ottaviano +6
Deploying DNNs on System-on-Chips (SoC) with multiple heterogeneous acceleration engines is challenging, and the majority of deployment frameworks cannot fully exploit heterogeneit…