1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Mohammadali Khodabandehlou, Bhaskar Krishnamachari
KV-cache compression reduces LLM inference memory by evicting context tokens, but when the evicted tokens contain answer-bearing evidence, the model may hallucinate instead of reco…