12 papers
DiffPrune: differentiable information throttling for token pruning in vision-language models
Landi He, Mingde Yao, Shawn Young +1
Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score that measures whether a token…
Decoupling semantics from vision: A framework for faithful visual-text compression evaluation
Yonghan Gao, Zehong Chen, Lijian Xu +3
Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image…
PathSelect: Sequential Token Selection for Whole Slide Pathology
Jingzhi Chen, Landi He, Zehong Chen +2
Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Existing approaches predominan…
Stepwise Token Selection for Efficient Multimodal Large Language Models
Landi He, Shawn Young, Lijian Xu
In multimodal large language models (MLLMs), inference cost is largely dominated by the visual token prefix rather than the language backbone, making token reduction a key factor f…
Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning
Jingzhi Chen, Landi He, Zhuo Chen +2
The processing of gigapixel whole slide images within vision language models faces a major difficulty due to an excessive number of visual tokens. Existing solutions typically rely…
A unified multi-task framework enables interpretable chest radiograph analysis
Lijian Xu, Ziyu Ni, Xinglong Liu +3
While multimodal deep learning has advanced medical imaging analysis, existing black-box systems \textcolor{black}{may remain confined to isolated tasks, often overlooking} the tru…