41 citations · 161 across the 37 of their papers we have counts for
5 papers · 2 filters
SceneGATE: Scene-Graph based co-Attention networks for TExt visual question answering
Feiqi Cao, Siwen Luo, Felipe Nunez +3
Most TextVQA approaches focus on the integration of objects, scene texts and question words by a simple transformer encoder. But this fails to capture the semantic relations betwee…
SG-Shuffle: Multi-aspect Shuffle Transformer for Scene Graph Generation
Anh Duc Bui, Soyeon Caren Han, Josiah Poon
Scene Graph Generation (SGG) serves a comprehensive representation of the images for human understanding as well as visual understanding tasks. Due to the long tail bias problem of…
Understanding Attention for Vision-and-Language Tasks
Feiqi Cao, Soyeon Caren Han, Siqu Long +2
Attention mechanism has been used as an important component across Vision-and-Language(VL) tasks in order to bridge the semantic gap between visual and textual features. While atte…
Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis
Siwen Luo, Yihao Ding, Siqu Long +2
Recognizing the layout of unstructured digital documents is crucial when parsing the documents into the structured, machine-readable format for downstream applications. Recent stud…
RoViST:Learning Robust Metrics for Visual Storytelling
Eileen Wang, Caren Han, Josiah Poon
Visual storytelling (VST) is the task of generating a story paragraph that describes a given image sequence. Most existing storytelling approaches have evaluated their models using…