3 papers
cs.CV2026
Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters
Gengluo Li, Shangpin Peng, Xingyu Wan +16
Vision Large Language Models (VLLMs) have achieved remarkable success in modern text-rich visual understanding. However, their perceptual robustness in the face of the continuous m…
cs.CV2025
The Role of Video Generation in Enhancing Data-Limited Action Understanding
Wei Li, Dezhao Luo, Dongbao Yang +3
Video action understanding tasks in real-world scenarios always suffer data limitations. In this paper, we address the data-limited action understanding problem by bridging data sc…
cs.CV2025
IPAD: Iterative, Parallel, and Diffusion-based Network for Scene Text Recognition
Xiaomeng Yang, Zhi Qiao, Yu Zhou
Nowadays, scene text recognition has attracted more and more attention due to its diverse applications. Most state-of-the-art methods adopt an encoder-decoder framework with the at…