445 citations · 447 across the 4 of their papers we have counts for
4 papers
TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Ya-Qi Yu, Minghui Liao, Jihao Wu +3
Multimodal Large Language Models (MLLMs) have shown impressive results on various multimodal tasks. However, most existing MLLMs are not well suited for document-oriented tasks, wh…
Sequential Visual and Semantic Consistency for Semi-supervised Text Recognition
Mingkun Yang, Biao Yang, Minghui Liao +2
Scene text recognition (STR) is a challenging task that requires large-scale annotated data for training. However, collecting and labeling real text images is expensive and time-co…
Class-Aware Mask-Guided Feature Refinement for Scene Text Recognition
Mingkun Yang, Biao Yang, Minghui Liao +2
Scene text recognition is a rapidly developing field that faces numerous challenges due to the complexity and diversity of scene text, including complex backgrounds, diverse fonts,…
TextBoxes: A Fast Text Detector with a Single Deep Neural Network
Minghui Liao, Baoguang Shi, Xiang Bai +2
This paper presents an end-to-end trainable fast scene text detector, named TextBoxes, which detects scene text with both high accuracy and efficiency in a single network forward p…