3 papers
cs.CV2026
Decoupling semantics from vision: A framework for faithful visual-text compression evaluation
Yonghan Gao, Zehong Chen, Lijian Xu +3
Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image…
cs.CV2026
ZeroSense:How Vision matters in Long Context Compression
Yonghan Gao, Zehong Chen, Lijian Xu +3
Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image…
cs.CV2025
Fewer Tokens, Greater Scaling: Self-Adaptive Visual Bases for Efficient and Expansive Representation Learning
Shawn Young, Xingyu Zeng, Lijian Xu
This paper investigates the fundamental relationship between model capacity and the minimal number of visual tokens required to preserve image semantics. Inspired by the Minimum De…