3 papers
cs.LG2025
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
Lion Mueller, Alberto Garcia-Ortiz, Ardalan Najafi +2
Integer AI inference significantly reduces computational complexity in embedded systems. Quantization-aware training (QAT) helps mitigate accuracy degradation associated with post-…
cs.AR2025
eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
Lennart Bamberg, Filippo Minnella, Roberto Bosio +5
Neural Processing Units (NPUs) are key to enabling efficient AI inference in resource-constrained edge environments. While peak tera operations per second (TOPS) is often used to g…
cs.AR2025
VUSA: Virtually Upscaled Systolic Array Architecture to Exploit Unstructured Sparsity in AI Acceleration
Shereef Helal, Alberto Garcia-Ortiz, Lennart Bamberg
Leveraging high degrees of unstructured sparsity is a promising approach to enhance the efficiency of deep neural network DNN accelerators - particularly important for emerging Edg…