2 papers
cs.LG2025
Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
Hui Zeng, Daming Zhao, Pengfei Yang +5
Generative reasoning with large language models (LLMs) often involves long decoding sequences, leading to substantial memory and latency overheads from accumulating key-value (KV)…
cs.CV2025
MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation
Hui Li, Pengfei Yang, Juanyang Chen +3
Knowledge distillation as an efficient knowledge transfer technique, has achieved remarkable success in unimodal scenarios. However, in cross-modal settings, conventional distillat…