2 papers
cs.AR2025
CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM
Qingyuan Liu, Liyan Chen, Yanning Yang +6
Attention-FC Disaggregated (AFD) LLM inference systems offload memory-bound Attention operations to memory-rich accelerators (e.g., CPUs, HBM-PIM) while retaining compute-bound Ful…
cs.AR2024
An Architectural Error Metric for CNN-Oriented Approximate Multipliers
Ao Liu, Jie Han, Qin Wang +2
As a potential alternative for implementing the large number of multiplications in convolutional neural networks (CNNs), approximate multipliers (AMs) promise both high hardware ef…