2 papers
cs.AR2025
BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language Models
Xiaomeng Han, Yuan Cheng, Jing Wang +6
Large language models (LLMs), with their billions of parameters, pose substantial challenges for deployment on edge devices, straining both memory capacity and computational resour…
cs.AR2025
NVR: Vector Runahead on NPUs for Sparse Memory Access
Hui Wang, Zhengpeng Zhao, Jing Wang +11
Deep Neural Networks are increasingly leveraging sparsity to reduce the scaling up of model parameter size. However, reducing wall-clock time through sparsity and pruning remains c…