Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
StepPrune: Adaptive Sequential Visual Token Selection across Multimodal Large Language Models
Hansen Zhang, Landi He, Mingde Yao +1
Visual prefixes account for a major portion of the per-layer computation in multimodal large language models (MLLMs), making visual-token pruning a direct approach to accelerating…
cs.CV2026
DiffPrune: differentiable information throttling for token pruning in vision-language models
Landi He, Mingde Yao, Shawn Young +1
Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score that measures whether a token…
cs.CV2026
Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models
Landi He, Mingde Yao, Shawn Young +1
Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. Existing methods typically rely on Gumbel-Softmax to appro…