1 paper
Siyuan He, Peiran Yan, Yandong He +2
The autoregressive decoding in LLMs is the major inference bottleneck due to the memory-intensive operations and limited hardware bandwidth. 3D-stacked architecture is a promising…