3 citations · 3 across the 2 of their papers we have counts for
1 paper · 1 filter
Tatsuya Kubo, Daichi Tokuda, Tomoya Nagatani +4
General matrix-vector multiplication (GeMV) remains a critical latency bottleneck in large language model (LLM) inference, even with quantized low-bit models. Processing-Using-DRAM…