3 papers
cs.AR2026
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
Chenhao Xue, Yukun Wang, An Guo +11
SRAM-based compute-in-memory (CIM) offers high computational density and energy efficiency for deep neural network (DNN) accelerators, but its limited capacity causes on/off-chip d…
cs.AR2026
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
Cong Li, Chenhao Xue, Yi Ren +11
Large language models (LLMs) exhibit memory-intensive behavior during decoding, making it a key bottleneck in LLM inference. To accelerate decoding execution, hybrid-bonding-based…
cs.AR2025
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
Yiqi Chen, Xiping Dong, Zhe Zhou +3
The Compute Express Link (CXL) technology facilitates the extension of CPU memory through byte-addressable SerDes links and cascaded switches, creating complex heterogeneous memory…