1 paper
Pingcheng Dong, Yonghao Tan, Xuejiao Liu +13
This work presents a 55nm speculative decoding-based LLM accelerator with bumping-based face-to-face ReRAM-on-logic stacking technology. It features a local rotation unit for outli…