6 citations · 7 across the 5 of their papers we have counts for
6 papers · 1 filter
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
Qian Kou, Xiaofeng Shi, Yulin Li +4
Multimodal Large Language Models (MLLMs) have demonstrated significant achievements in general visual question answering (VQA) tasks. However, they remain brittle on mechanical eng…
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
Pengyuan Lyu, Yulin Li, Hao Zhou +8
Text-rich images have significant and extensive value, deeply integrated into various aspects of human life. Notably, both visual cues and linguistic symbols in text-rich images pl…
Collaborative Position Reasoning Network for Referring Image Segmentation
Jianjian Cao, Beiya Dai, Yulin Li +2
Given an image and a natural language expression as input, the goal of referring image segmentation is to segment the foreground masks of the entities referred by the expression. E…
MataDoc: Margin and Text Aware Document Dewarping for Arbitrary Boundary
Beiya Dai, Xing li, Qunyi Xie +5
Document dewarping from a distorted camera-captured image is of great value for OCR and document understanding. The document boundary plays an important role which is more evident…
Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document Understanding
Mingliang Zhai, Yulin Li, Xiameng Qin +6
Transformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the…
StrucTexT: Structured Text Understanding with Multi-Modal Transformers
Yulin Li, Yuxi Qian, Yuchen Yu +7
Structured text understanding on Visually Rich Documents (VRDs) is a crucial part of Document Intelligence. Due to the complexity of content and layout in VRDs, structured text und…