most citedPDFVQA: A New Dataset for Real-World VQA on PDF Documents

2 citations · 3 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CL2024

Multimodal Commonsense Knowledge Distillation for Visual Question Answering

Shuo Yang, Siwen Luo, Soyeon Caren Han

Existing Multimodal Large Language Models (MLLMs) and Visual Language Pretrained Models (VLPMs) have shown remarkable performances in the general Visual Question Answering (VQA). H…

cs.CL20241 cited

MSG-Chart: Multimodal Scene Graph for ChartQA

Yue Dai, Soyeon Caren Han, Wei Liu

Automatic Chart Question Answering (ChartQA) is challenging due to the complex distribution of chart elements with patterns of the underlying data not explicitly displayed in chart…

cs.CL2024

3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection

Thye Shan Ng, Feiqi Cao, Soyeon Caren Han

Esports has rapidly emerged as a global phenomenon with an ever-expanding audience via platforms, like YouTube. Due to the inherent complexity nature of the game, it is challenging…

cs.CV2024

PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering

Yihao Ding, Kaixuan Ren, Jiabin Huang +2

Document Question Answering (QA) presents a challenge in understanding visually-rich documents (VRD), particularly those dominated by lengthy textual content like research journal…

cs.CV2024

SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling

Eileen Wang, Soyeon Caren Han, Josiah Poon

Visual storytelling aims to automatically generate a coherent story based on a given image sequence. Unlike tasks like image captioning, visual stories should contain factual descr…

cs.CL2023

Re-Temp: Relation-Aware Temporal Representation Learning for Temporal Knowledge Graph Completion

Kunze Wang, Soyeon Caren Han, Josiah Poon

Temporal Knowledge Graph Completion (TKGC) under the extrapolation setting aims to predict the missing entity from a fact in the future, posing a challenge that aligns more closely…