2 citations · 2 across the 1 of their papers we have counts for
1 paper
Taesu Kim, Jongho Lee, Daehyun Ahn +4
We introduce QUICK, a group of novel optimized CUDA kernels for the efficient inference of quantized Large Language Models (LLMs). QUICK addresses the shared memory bank-conflict p…