12 citations · 12 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 12 cited
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Yuliang Liu, Biao Yang, Qiang Liu +4
We present TextMonkey, a large multimodal model (LMM) tailored for text-centric tasks. Our approach introduces enhancement across several dimensions: By adopting Shifted Window Att…
cs.CV2024
Sequential Visual and Semantic Consistency for Semi-supervised Text Recognition
Mingkun Yang, Biao Yang, Minghui Liao +2
Scene text recognition (STR) is a challenging task that requires large-scale annotated data for training. However, collecting and labeling real text images is expensive and time-co…