3 papers
cs.CV2026
Decoupling semantics from vision: A framework for faithful visual-text compression evaluation
Yonghan Gao, Zehong Chen, Lijian Xu +3
Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image…
cs.CV2026
PathSelect: Sequential Token Selection for Whole Slide Pathology
Jingzhi Chen, Landi He, Zehong Chen +2
Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Existing approaches predominan…
cs.CV2026
Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning
Jingzhi Chen, Landi He, Zhuo Chen +2
The processing of gigapixel whole slide images within vision language models faces a major difficulty due to an excessive number of visual tokens. Existing solutions typically rely…