2 citations · 3 across the 10 of their papers we have counts for
10 papers
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
Shuo Yang, Siwen Luo, Soyeon Caren Han
Existing Multimodal Large Language Models (MLLMs) and Visual Language Pretrained Models (VLPMs) have shown remarkable performances in the general Visual Question Answering (VQA). H…
MSG-Chart: Multimodal Scene Graph for ChartQA
Yue Dai, Soyeon Caren Han, Wei Liu
Automatic Chart Question Answering (ChartQA) is challenging due to the complex distribution of chart elements with patterns of the underlying data not explicitly displayed in chart…
3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection
Thye Shan Ng, Feiqi Cao, Soyeon Caren Han
Esports has rapidly emerged as a global phenomenon with an ever-expanding audience via platforms, like YouTube. Due to the inherent complexity nature of the game, it is challenging…
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
Yihao Ding, Kaixuan Ren, Jiabin Huang +2
Document Question Answering (QA) presents a challenge in understanding visually-rich documents (VRD), particularly those dominated by lengthy textual content like research journal…
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
Eileen Wang, Soyeon Caren Han, Josiah Poon
Visual storytelling aims to automatically generate a coherent story based on a given image sequence. Unlike tasks like image captioning, visual stories should contain factual descr…
Re-Temp: Relation-Aware Temporal Representation Learning for Temporal Knowledge Graph Completion
Kunze Wang, Soyeon Caren Han, Josiah Poon
Temporal Knowledge Graph Completion (TKGC) under the extrapolation setting aims to predict the missing entity from a fact in the future, posing a challenge that aligns more closely…