1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.AR2025
CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM
Qingyuan Liu, Liyan Chen, Haocheng Wang +6
Attention-FC Disaggregated (AFD) LLM inference systems offload memory-bound Attention operations to memory-rich accelerators (e.g., CPUs, HBM-PIM) while retaining compute-bound Ful…
cs.AR2024★ 1 cited
An Architectural Error Metric for CNN-Oriented Approximate Multipliers
Ao Liu, Jie Han, Qin Wang +2
As a potential alternative for implementing the large number of multiplications in convolutional neural networks (CNNs), approximate multipliers (AMs) promise both high hardware ef…