3 papers
cs.CV2026
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
Xumin Yu, Zuyan Liu, Zhenyu Yang +5
A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, representing images as discrete s…
cs.CV2026
Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters
Gengluo Li, Shangpin Peng, Xingyu Wan +16
Vision Large Language Models (VLLMs) have achieved remarkable success in modern text-rich visual understanding. However, their perceptual robustness in the face of the continuous m…
cs.CV2023
Segment and Caption Anything
Xiaoke Huang, Jianfeng Wang, Yansong Tang +5
We propose a method to efficiently equip the Segment Anything Model (SAM) with the ability to generate regional captions. SAM presents strong generalizability to segment anything w…