10 citations · 10 across the 1 of their papers we have counts for
3 papers
cs.AR2025
A Digital SRAM-Based Compute-In-Memory Macro for Weight-Stationary Dynamic Matrix Multiplication in Transformer Attention Score Computation
Jianyi Yu, Tengxiao Wang, Yuxuan Wang +6
Compute-in-memory (CIM) techniques are widely employed in energy-efficient artificial intelligent (AI) processors. They alleviate power and latency bottlenecks caused by extensive…
cs.AR2024
COMET: Towards Partical W4A4KV4 LLMs Serving
Lian Liu, Haimeng Ren, Long Cheng +6
Quantization is a widely-used compression technology to reduce the overhead of serving large language models (LLMs) on terminal devices and in cloud data centers. However, prevalen…
cs.AR2024★ 10 cited
CIM-MLC: A Multi-level Compilation Stack for Computing-In-Memory Accelerators
Songyun Qu, Shixin Zhao, Bing Li +4
In recent years, various computing-in-memory (CIM) processors have been presented, showing superior performance over traditional architectures. To unleash the potential of various…