papers

Publications (8)

cs.AR2021

MPU: Towards Bandwidth-abundant SIMT Processor via Near-bank Computing

Xinfeng Xie, Peng Gu, Yufei Ding +3

With the growing number of data-intensive workloads, GPU, which is the state-of-the-art single-instruction-multiple-thread (SIMT) processor, is hindered by the memory bandwidth wal…

cs.DC2022

Accelerating CPU-Based Sparse General Matrix Multiplication With Binary Row Merging

Zhaoyang Du, Yijin Guan, Tianchan Guan +3

Sparse general matrix multiplication (SpGEMM) is a fundamental building block for many real-world applications. Since SpGEMM is a well-known memory-bounded application with vast an…

cs.DC2022

OpSparse: a Highly Optimized Framework for Sparse General Matrix Multiplication on GPUs

Zhaoyang Du, Yijin Guan, Tianchan Guan +4

Sparse general matrix multiplication (SpGEMM) is an important and expensive computation primitive in many real-world applications. Due to SpGEMM's inherent irregularity and the vas…

cs.AR2025

Enabling Efficient Transaction Processing on CXL-Based Memory Sharing

Zhao Wang, Yiqi Chen, Cong Li +5

Transaction processing systems are the crux for modern data-center applications, yet current multi-node systems are slow due to network overheads. This paper advocates for Compute…

cs.AR2026

CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation

Yuanpeng Zhang, YuXuan Wu, Yitong Xiao +6

The paper introduces CODA, a hardware-software co-designed architecture that separates compute and cache operations for edge video diffusion models, using near‑memory processing to…

#edge computing#video diffusion models#cross‑timestep caching#near‑memory processing
cs.PF2025

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing

Haochen Huang, Shuzhang Zhong, Zhe Zhang +5

Large Language Models (LLMs) with Mixture-of-Expert (MoE) architectures achieve superior model performance with reduced computation costs, but at the cost of high memory capacity a…

cs.CV2024

Unicorn: Unified Neural Image Compression with One Number Reconstruction

Qi Zheng, Haozhi Wang, Zihao Liu +8

Prevalent lossy image compression schemes can be divided into: 1) explicit image compression (EIC), including traditional standards and neural end-to-end algorithms; 2) implicit im…

cs.DC2022

Predicting the Output Structure of Sparse Matrix Multiplication with Sampled Compression Ratio

Zhaoyang Du, Yijin Guan, Tianchan Guan +7

Sparse general matrix multiplication (SpGEMM) is a fundamental building block in numerous scientific applications. One critical task of SpGEMM is to compute or predict the structur…