81 citations · 82 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
Yuting Li, Lai Wei, Kaipeng Zheng +6
Despite the rapid progress of multimodal large language models (MLLMs), they have largely overlooked the importance of visual processing. In a simple yet revealing experiment, we i…
cs.CV2024★ 81 cited
HTR-VT: Handwritten Text Recognition with Vision Transformer
Yuting Li, Dexiong Chen, Tinglong Tang +1
We explore the application of Vision Transformer (ViT) for handwritten text recognition. The limited availability of labeled data in this domain poses challenges for achieving high…