2 citations · 3 across the 5 of their papers we have counts for
8 papers
HRVDA: High-Resolution Visual Document Assistant
Chaohu Liu, Kun Yin, Haoyu Cao +6
Leveraging vast training data, multimodal large language models (MLLMs) have demonstrated formidable general visual comprehension capabilities and achieved remarkable performance a…
Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region Concentration
Haoyu Cao, Changcun Bao, Chaohu Liu +6
We propose a novel end-to-end document understanding model called SeRum (SElective Region Understanding Model) for extracting meaningful information from document images, including…
Looking and Listening: Audio Guided Text Recognition
Wenwen Yu, Mingyu Liu, Biao Yang +5
Text recognition in the wild is a long-standing problem in computer vision. Driven by end-to-end deep learning, recent studies suggest vision and language processing are effective…
Turning a CLIP Model into a Scene Text Detector
Wenwen Yu, Yuliang Liu, Wei Hua +3
The recent large-scale Contrastive Language-Image Pretraining (CLIP) model has shown great potential in various downstream tasks via leveraging the pretrained vision and language k…
Grab What You Need: Rethinking Complex Table Structure Recognition with Flexible Components Deliberation
Hao Liu, Xin Li, Mingming Gong +5
Recently, Table Structure Recognition (TSR) task, aiming at identifying table structure into machine readable formats, has received increasing interest in the community. While impr…
TaCo: Textual Attribute Recognition via Contrastive Learning
Chang Nie, Yiqing Hu, Yanqiu Qu +3
As textual attributes like font are core design elements of document format and page style, automatic attributes recognition favor comprehensive practical applications. Existing ap…