activity
20202026
most citedDoc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis

8 citations · 8 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2025

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends

Yihao Ding, Siwen Luo, Yue Dai +6

Visually Rich Document Understanding (VRDU) has become a pivotal area of research, driven by the need to automatically interpret documents that contain intricate visual, textual, a…

cs.CV2024

PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering

Yihao Ding, Kaixuan Ren, Jiabin Huang +2

Document Question Answering (QA) presents a challenge in understanding visually-rich documents (VRD), particularly those dominated by lengthy textual content like research journal…

cs.CV20232 cited

PDFVQA: A New Dataset for Real-World VQA on PDF Documents

Yihao Ding, Siwen Luo, Hyunsuk Chung +1

Document-based Visual Question Answering examines the document understanding of document images in conditions of natural language questions. We proposed a new document-based VQA da…

cs.CV2022

PiggyBack: Pretrained Visual Question Answering Environment for Backing up Non-deep Learning Professionals

Zhihao Zhang, Siwen Luo, Junyi Chen +4

We propose a PiggyBack, a Visual Question Answering platform that allows users to apply the state-of-the-art visual-language pretrained models easily. The PiggyBack supports the fu…

cs.CV20228 cited

Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis

Siwen Luo, Yihao Ding, Siqu Long +2

Recognizing the layout of unstructured digital documents is crucial when parsing the documents into the structured, machine-readable format for downstream applications. Recent stud…

cs.CV2020

VICTR: Visual Information Captured Text Representation for Text-to-Image Multimodal Tasks

Soyeon Caren Han, Siqu Long, Siwen Luo +2

Text-to-image multimodal tasks, generating/retrieving an image from a given text description, are extremely challenging tasks since raw text descriptions cover quite limited inform…