3 papers
cs.AI2026
GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving
Qiankun Ma, Yanjiang Zhou, Zinan Xiong +5
Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-re…
cs.CV2025
Training-free Token Reduction for Vision Mamba
Qiankun Ma, Ziyao Zhang, Chi Su +4
Vision Mamba has emerged as a strong competitor to Vision Transformers (ViTs) due to its ability to efficiently capture long-range dependencies with linear computational complexity…
cs.CE2024
BrainMass: Advancing Brain Network Analysis for Diagnosis with Large-scale Self-Supervised Learning
Yanwu Yang, Chenfei Ye, Guinan Su +6
Foundation models pretrained on large-scale datasets via self-supervised learning demonstrate exceptional versatility across various tasks. Due to the heterogeneity and hard-to-col…