1 paper
Ye Lin, Chao Fang, Xiaoyong Song +4
Edge deployment of low-batch large language models (LLMs) faces critical memory bandwidth bottlenecks when executing memory-intensive general matrix-vector multiplications (GEMV) o…