Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compression
Runchao Li, Yao Fu, Mu Sheng +3
The efficacy of Large Language Models (LLMs) in long-context tasks is often hampered by the substantial memory footprint and computational demands of the Key-Value (KV) cache. Curr…
cs.CL2024
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models
Yao Fu, Yin Yu, Xiaotian Han +4
Knowledge distillation (KD) has become a widely adopted approach for compressing large language models (LLMs) to reduce computational costs and memory footprints. However, the avai…