3 papers
cs.AR2025
An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator Systems
Ryan Albert Antonio, Joren Dumoulin, Xiaoling Yi +4
Heterogeneous accelerator-centric compute clusters are emerging as efficient solutions for diverse AI workloads. However, current integration strategies often compromise data movem…
cs.AR2024
OpenGeMM: A High-Utilization GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory Coupling
Xiaoling Yi, Ryan Antonio, Joren Dumoulin +4
Deep neural networks (DNNs) face significant challenges when deployed on resource-constrained extreme edge devices due to their computational and data-intensive nature. While stand…
cs.PL2024
HTVM: Efficient Neural Network Deployment On Heterogeneous TinyML Platforms
Josse Van Delm, Maarten Vandersteegen, Alessio Burrello +5
Optimal deployment of deep neural networks (DNNs) on state-of-the-art Systems-on-Chips (SoCs) is crucial for tiny machine learning (TinyML) at the edge. The complexity of these SoC…