3 papers
cs.AI2026
GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving
Qiankun Ma, Yanjiang Zhou, Zinan Xiong +5
Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-re…
cs.CV2026
ApET: Approximation-Error Guided Token Compression for Efficient VLMs
Qiankun Ma, Ziyao Zhang, Haofei Wang +3
Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead an…
cs.CV2025
Training-free Token Reduction for Vision Mamba
Qiankun Ma, Ziyao Zhang, Chi Su +4
Vision Mamba has emerged as a strong competitor to Vision Transformers (ViTs) due to its ability to efficiently capture long-range dependencies with linear computational complexity…