10 citations · 10 across the 5 of their papers we have counts for
Showing cs.ARShow all
3 papers · 1 filter
cs.AR2025
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
Shixin Zhao, Yuming Li, Bing Li +4
Computing-in-memory (CIM) architectures demonstrate superior performance over traditional architectures. To unleash the potential of CIM accelerators, many compilation methods have…
cs.AR2025
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
Lian Liu, Shixin Zhao, Bing Li +6
The billion-scale Large Language Models (LLMs) need deployment on expensive server-grade GPUs with large-storage HBMs and abundant computation capability. As LLM-assisted services…
cs.AR2024★ 10 cited
CIM-MLC: A Multi-level Compilation Stack for Computing-In-Memory Accelerators
Songyun Qu, Shixin Zhao, Bing Li +4
In recent years, various computing-in-memory (CIM) processors have been presented, showing superior performance over traditional architectures. To unleash the potential of various…