3 papers
cs.CV2026
Masked Next-Scale Prediction for Self-supervised Scene Text Recognition
Zhuohao Chen, Zeng Li, Yifei Zhang +2
Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine-grained character strokes. Training such models relies on large amounts of annota…
cs.CV2026
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
Xuehai Bai, Yang Shi, Yi-Fan Zhang +7
Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fai…
cs.CV2025
Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
Yifei Zhang, Chang Liu, Jin Wei +4
Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while t…