collaborators

5 papers

cs.CV2026

Decoupling semantics from vision: A framework for faithful visual-text compression evaluation

Yonghan Gao, Zehong Chen, Lijian Xu +3

Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image…

cs.CV2026

PathSelect: Sequential Token Selection for Whole Slide Pathology

Jingzhi Chen, Landi He, Zehong Chen +2

Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Existing approaches predominan…

cs.CV2026

Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning

Jingzhi Chen, Landi He, Zhuo Chen +2

The processing of gigapixel whole slide images within vision language models faces a major difficulty due to an excessive number of visual tokens. Existing solutions typically rely…

cs.CV2026

From Snapshots to Symphonies: The Evolution of Protein Prediction from Static Structures to Generative Dynamics and Multimodal Interactions

Jingzhi Chen, Lijian Xu

The protein folding problem has been fundamentally transformed by artificial intelligence, evolving from static structure prediction toward the modeling of dynamic conformational e…

cs.CV2026

ZeroSense:How Vision matters in Long Context Compression

Yonghan Gao, Zehong Chen, Lijian Xu +3

Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image…