11 citations · 12 across the 2 of their papers we have counts for
Showing cs.ARShow all
2 papers · 1 filter
cs.AR2026★ 1 cited
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
Pingcheng Dong, Yonghao Tan, Xuejiao Liu +13
This work presents a 55nm speculative decoding-based LLM accelerator with bumping-based face-to-face ReRAM-on-logic stacking technology. It features a local rotation unit for outli…
cs.AR2025
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
Yonghao Tan, Pingcheng Dong, Yongkun Wu +8
DNN accelerators, significantly advanced by model compression and specialized dataflow techniques, have marked considerable progress. However, the frequent access of high-precision…