8 papers
DiffPrune: differentiable information throttling for token pruning in vision-language models
Landi He, Mingde Yao, Shawn Young +1
Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score that measures whether a token…
Stepwise Token Selection for Efficient Multimodal Large Language Models
Landi He, Shawn Young, Lijian Xu
In multimodal large language models (MLLMs), inference cost is largely dominated by the visual token prefix rather than the language backbone, making token reduction a key factor f…
Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning
Jingzhi Chen, Landi He, Zhuo Chen +2
The processing of gigapixel whole slide images within vision language models faces a major difficulty due to an excessive number of visual tokens. Existing solutions typically rely…
Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models
Landi He, Mingde Yao, Shawn Young +1
Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. Existing methods typically rely on Gumbel-Softmax to appro…
XrayClaw: Cooperative-Competitive Multi-Agent Alignment for Trustworthy Chest X-ray Diagnosis
Shawn Young, Lijian Xu
Chest X-ray (CXR) interpretation is a fundamental yet complex clinical task that increasingly relies on artificial intelligence for automation. However, traditional monolithic mode…
Efficient Chest X-ray Representation Learning via Semantic-Partitioned Contrastive Learning
Wangyu Feng, Shawn Young, Lijian Xu
Self-supervised learning (SSL) has emerged as a powerful paradigm for Chest X-ray (CXR) analysis under limited annotations. Yet, existing SSL strategies remain suboptimal for medic…