2 papers
cs.LG2025
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
Lion Mueller, Alberto Garcia-Ortiz, Ardalan Najafi +2
Integer AI inference significantly reduces computational complexity in embedded systems. Quantization-aware training (QAT) helps mitigate accuracy degradation associated with post-…
cs.AR2025
eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
Lennart Bamberg, Filippo Minnella, Roberto Bosio +5
Neural Processing Units (NPUs) are key to enabling efficient AI inference in resource-constrained edge environments. While peak tera operations per second (TOPS) is often used to g…