4 papers
Decoupling semantics from vision: A framework for faithful visual-text compression evaluation
Yonghan Gao, Zehong Chen, Lijian Xu +3
Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image…
PathSelect: Sequential Token Selection for Whole Slide Pathology
Jingzhi Chen, Landi He, Zehong Chen +2
Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Existing approaches predominan…
Multimodal Model for Computational Pathology:Representation Learning and Image Compression
Peihang Wu, Zehong Chen, Lijian Xu
Whole slide imaging (WSI) has transformed digital pathology by enabling computational analysis of gigapixel histopathology images. Recent foundation model advances have accelerated…
ZeroSense:How Vision matters in Long Context Compression
Yonghan Gao, Zehong Chen, Lijian Xu +3
Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image…