2 papers
cs.AR2026
PENDA: An Efficient Processing Element via Norm-of-Difference for Deep Learning Accelerators
Kai-Chieh Hsu, Tian-Sheuan Chang
Inner product computation dominates the computational cost of deep learning models; thus, accelerating this primitive is key to improving hardware efficiency. However, most existin…
cs.AR2025
A Low-Power Sparse Deep Learning Accelerator with Optimized Data Reuse
Kai-Chieh Hsu, Tian-Sheuan Chang
Sparse deep learning has reduced computation significantly, but its irregular non-zero data distribution complicates the data flow and hinders data reuse, increasing on-chip SRAM a…