1 paper · 1 filter
Shiwei Gao, Youmin Chen, Jiwu Shu
The growing complexity of LLM usage today, e.g., multi-round conversation and retrieval-augmented generation (RAG), makes contextual states (i.e., KV cache) reusable across user re…