1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Ye Lin, Chao Fang, Xiaoyong Song +4
Edge deployment of low-batch large language models (LLMs) faces critical memory bandwidth bottlenecks when executing memory-intensive general matrix-vector multiplications (GEMV) o…