1 paper
Siyan Zhao, Daniel Israel, Guy Van den Broeck +1
During inference for transformer-based large language models (LLM), prefilling is the computation of the key-value (KV) cache for input tokens in the prompt prior to autoregressive…