paper

Quantization Effects of Artificial Neural Networks for Embedded Edge-Computing Applications

arXiv:2511.05479

Abstract

This paper examines the use of Quantized Neural Networks (QNNs) for two resource-constrained scientific applications: automated calibration of semi- conductor quantum bits (qubits) and scientific particle detectors. We evaluate the trade-offs between Post-Training Quantization (PTQ), Quantization-Aware Train- ing (QAT), and ultra-low-bit Binary Neural Networks (BNNs) with respect to la- tency and resource usage. Our results demonstrate that PTQ and QAT are easily implementable solutions achieving a four-fold reduction in memory usage for U- shaped CNN (U-Net) architectures, whereas BNNs are better suitable for use cases with challenging latency requirements. For the training of non-differentiable custom BNNs , we propose a novel, hardware-constrained learning approach using Genetic Algorithms (GAs). We showcase a Look-Up Table (LUT)-based BNN architecture suitable for direct conversion to Very High-Speed Integrated Circuit Hardware De- scription Language (VHDL) via the HCL4BNN framework. This method achieves nanosecond-scale inference latencies at 10 ns to 15 ns without requiring specialized Digital Signal Processor (DSP) or Block RAM (BRAM) resources.

deRSE26 proceedings pre-print

Quantization Effects of Artificial Neural Networks for Embedded Edge-Computing Applications · wovepaper