Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM
Shuvendu Roy, Mengyao Zhai, Hossein Hajimirsadeghi +1
Large language models (LLMs) excel at complex tasks like question answering and summarization, thanks to their ability to handle long-context inputs. However, deploying LLMs is cos…
cs.CL2025
Task-agnostic Prompt Compression with Context-aware Sentence Embedding and Reward-guided Task Descriptor
Barys Liskavets, Shuvendu Roy, Maxim Ushakov +3
The rise of Large Language Models (LLMs) has led to significant interest in prompt compression, a technique aimed at reducing the length of input prompts while preserving critical…
cs.CL2024
Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference
Barys Liskavets, Maxim Ushakov, Shuvendu Roy +3
Large language models (LLMs) have triggered a new stream of research focusing on compressing the context length to reduce the computational cost while ensuring the retention of hel…