collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2024

WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation

Zirui Shao, Feiyu Gao, Hangdi Xing +5

In the era of content creation revolution propelled by advancements in generative models, the field of web design remains unexplored despite its critical role in modern digital com…

cs.CV2024

ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction Data

Yufan Shen, Chuwei Luo, Zhaoqing Zhu +5

Recently, large language models (LLMs) and multimodal large language models (MLLMs) have demonstrated promising results on document visual question answering (VQA) task, particular…

cs.CV20244 cited

LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding

Chuwei Luo, Yufan Shen, Zhaoqing Zhu +3

Recently, leveraging large language models (LLMs) or multimodal large language models (MLLMs) for document understanding has been proven very promising. However, previous works tha…

cs.CV2024

LORE++: Logical Location Regression Network for Table Structure Recognition with Pre-training

Rujiao Long, Hangdi Xing, Zhibo Yang +4

Table structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations o…

cs.CV20232 cited

Vision Grid Transformer for Document Layout Analysis

Cheng Da, Chuwei Luo, Qi Zheng +1

Document pre-trained models and grid-based models have proven to be very effective on various tasks in Document AI. However, for the document layout analysis (DLA) task, existing d…

cs.CV20232 cited

LISTER: Neighbor Decoding for Length-Insensitive Scene Text Recognition

Changxu Cheng, Peng Wang, Cheng Da +2

The diversity in length constitutes a significant characteristic of text. Due to the long-tail distribution of text lengths, most existing methods for scene text recognition (STR)…