Publications (8)
MPU: Towards Bandwidth-abundant SIMT Processor via Near-bank Computing
Xinfeng Xie, Peng Gu, Yufei Ding +3
With the growing number of data-intensive workloads, GPU, which is the state-of-the-art single-instruction-multiple-thread (SIMT) processor, is hindered by the memory bandwidth wal…
Accelerating CPU-Based Sparse General Matrix Multiplication With Binary Row Merging
Zhaoyang Du, Yijin Guan, Tianchan Guan +3
Sparse general matrix multiplication (SpGEMM) is a fundamental building block for many real-world applications. Since SpGEMM is a well-known memory-bounded application with vast an…
OpSparse: a Highly Optimized Framework for Sparse General Matrix Multiplication on GPUs
Zhaoyang Du, Yijin Guan, Tianchan Guan +4
Sparse general matrix multiplication (SpGEMM) is an important and expensive computation primitive in many real-world applications. Due to SpGEMM's inherent irregularity and the vas…
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
Zhao Wang, Yiqi Chen, Cong Li +5
Transaction processing systems are the crux for modern data-center applications, yet current multi-node systems are slow due to network overheads. This paper advocates for Compute…
CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation
Yuanpeng Zhang, YuXuan Wu, Yitong Xiao +6
The paper introduces CODA, a hardware-software co-designed architecture that separates compute and cache operations for edge video diffusion models, using near‑memory processing to…
HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
Haochen Huang, Shuzhang Zhong, Zhe Zhang +5
Large Language Models (LLMs) with Mixture-of-Expert (MoE) architectures achieve superior model performance with reduced computation costs, but at the cost of high memory capacity a…
Unicorn: Unified Neural Image Compression with One Number Reconstruction
Qi Zheng, Haozhi Wang, Zihao Liu +8
Prevalent lossy image compression schemes can be divided into: 1) explicit image compression (EIC), including traditional standards and neural end-to-end algorithms; 2) implicit im…
Predicting the Output Structure of Sparse Matrix Multiplication with Sampled Compression Ratio
Zhaoyang Du, Yijin Guan, Tianchan Guan +7
Sparse general matrix multiplication (SpGEMM) is a fundamental building block in numerous scientific applications. One critical task of SpGEMM is to compute or predict the structur…