Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
Shunlong Wu, Hai Lin, Shaoshen Chen +5
Existing KV cache compression methods generally operate on discrete tokens or non-semantic chunks. However, such approaches often lead to semantic fragmentation, where linguistical…
cs.CL2025
GMSA: Enhancing Context Compression via Group Merging and Layer Semantic Alignment
Jiwei Tang, Zhicheng Zhang, Shunlong Wu +8
Large Language Models (LLMs) have achieved remarkable performance across a wide range of Natural Language Processing (NLP) tasks. However, in long-context scenarios, they face two…