collaborators

5 papers

cs.CV2026

DiffPrune: differentiable information throttling for token pruning in vision-language models

Landi He, Mingde Yao, Shawn Young +1

Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score that measures whether a token…

cs.CV2026

PathSelect: Sequential Token Selection for Whole Slide Pathology

Jingzhi Chen, Landi He, Zehong Chen +2

Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Existing approaches predominan…

cs.CV2026

Stepwise Token Selection for Efficient Multimodal Large Language Models

Landi He, Shawn Young, Lijian Xu

In multimodal large language models (MLLMs), inference cost is largely dominated by the visual token prefix rather than the language backbone, making token reduction a key factor f…

cs.CV2026

Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning

Jingzhi Chen, Landi He, Zhuo Chen +2

The processing of gigapixel whole slide images within vision language models faces a major difficulty due to an excessive number of visual tokens. Existing solutions typically rely…

cs.CV2026

Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models

Landi He, Mingde Yao, Shawn Young +1

Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. Existing methods typically rely on Gumbel-Softmax to appro…