Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
Da Ma, Lu Chen, Situo Zhang +8
The rapid expansion of context window sizes in Large Language Models~(LLMs) has enabled them to tackle increasingly complex tasks involving lengthy documents. However, this progres…
cs.CL2025
Alignment for Efficient Tool Calling of Large Language Models
Hongshen Xu, Zihan Wang, Zichen Zhu +4
Recent advancements in tool learning have enabled large language models (LLMs) to integrate external tools, enhancing their task performance by expanding their knowledge boundaries…