floating-point formats 1hardware-aware training 1large language models 1mixed-precision quantization 1model compression 1
From the 1 of 14 linked papers with an AI index.
Showing cs.ARShow all
2 papers · 1 filter
cs.AR2026
DataGuard: Guaranteeing Private Training in Systolic-array Based Accelerators
Pawan Kumar Sanjaya, Christina Giannoula, Nikhil Shreekumar +6
Differential privacy (DP) and federated learning (FL) have emerged as important privacy-preserving approaches when using sensitive data to train machine learning (ML) models. FL en…
cs.AR2025
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
Yaman Umuroglu, Christoph Berganski, Felix Jentzsch +8
While neural network quantization effectively reduces the cost of matrix multiplications, aggressive quantization can expose non-matrix-multiply operations as significant performan…