3 papers
cs.AR2026
TetrisG-SDK: Efficient Convolutional Layer Mapping with Adaptive Windows and Grouped Convolutions for Fast In-Memory Computing
Ke Dong, Kejie Huang, Tao Luo +1
Shifted-and-Duplicated-Kernel (SDK) mapping has emerged as an effective strategy to accelerate convolutional layers on compute-in-memory (CIM) hardware. However, existing SDK varia…
cs.LG2026
DAPA: Distribution Aware Piecewise Activation Functions for On-Device Transformer Inference and Training
Maoyang Xiang, Bo Wang
Non-linear activation functions play a pivotal role in on-device inference and training, as they not only consume substantial hardware resources but also impose a significant impac…
cs.LG2025
Coflex: Enhancing HW-NAS with Sparse Gaussian Processes for Efficient and Scalable DNN Accelerator Design
Yinhui Ma, Tomomasa Yamasaki, Zhehui Wang +2
Hardware-Aware Neural Architecture Search (HW-NAS) is an efficient approach to automatically co-optimizing neural network performance and hardware energy efficiency, making it part…