4 papers · 1 filter
A High-Level Compiler Integration Approach for Deep Learning Accelerators Supporting Abstraction and Optimization
Samira Ahmadifarsani, Daniel Mueller-Gritschneder, Ulf Schlichtmann
The growing adoption of domain-specific architectures in edge computing platforms for deep learning has highlighted the efficiency of hardware accelerators. However, integrating cu…
TinyProp -- Adaptive Sparse Backpropagation for Efficient TinyML On-device Learning
Marcus Rüb, Daniel Maier, Daniel Mueller-Gritschneder +1
Training deep neural networks using backpropagation is very memory and computationally intensive. This makes it difficult to run on-device learning or fine-tune neural networks on…
MLonMCU: TinyML Benchmarking with Fast Retargeting
Philipp van Kempen, Rafael Stahl, Daniel Mueller-Gritschneder +1
While there exist many ways to deploy machine learning models on microcontrollers, it is non-trivial to choose the optimal combination of frameworks and targets for a given applica…
Fused Depthwise Tiling for Memory Optimization in TinyML Deep Neural Network Inference
Rafael Stahl, Daniel Mueller-Gritschneder, Ulf Schlichtmann
Memory optimization for deep neural network (DNN) inference gains high relevance with the emergence of TinyML, which refers to the deployment of DNN inference tasks on tiny, low-po…