41 citations · 111 across the 27 of their papers we have counts for
9 papers · 1 filter
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
Yihao Ding, Soyeon Caren Han, Yan Li +1
Visually Rich Document Understanding (VRDU) has emerged as a critical field in document intelligence, enabling automated extraction of key information from complex documents across…
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
Eileen Wang, Caren Han, Josiah Poon
Video Paragraph Captioning (VPC) aims to generate paragraph captions that summarises key events within a video. Despite recent advancements, challenges persist, notably in effectiv…
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
Eileen Wang, Soyeon Caren Han, Josiah Poon
Visual storytelling aims to automatically generate a coherent story based on a given image sequence. Unlike tasks like image captioning, visual stories should contain factual descr…
SG-Shuffle: Multi-aspect Shuffle Transformer for Scene Graph Generation
Anh Duc Bui, Soyeon Caren Han, Josiah Poon
Scene Graph Generation (SGG) serves a comprehensive representation of the images for human understanding as well as visual understanding tasks. Due to the long tail bias problem of…
Understanding Attention for Vision-and-Language Tasks
Feiqi Cao, Soyeon Caren Han, Siqu Long +2
Attention mechanism has been used as an important component across Vision-and-Language(VL) tasks in order to bridge the semantic gap between visual and textual features. While atte…
Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis
Siwen Luo, Yihao Ding, Siqu Long +2
Recognizing the layout of unstructured digital documents is crucial when parsing the documents into the structured, machine-readable format for downstream applications. Recent stud…