41 citations · 111 across the 16 of their papers we have counts for
20 papers
PiggyBack: Pretrained Visual Question Answering Environment for Backing up Non-deep Learning Professionals
Zhihao Zhang, Siwen Luo, Junyi Chen +4
We propose a PiggyBack, a Visual Question Answering platform that allows users to apply the state-of-the-art visual-language pretrained models easily. The PiggyBack supports the fu…
SG-Shuffle: Multi-aspect Shuffle Transformer for Scene Graph Generation
Anh Duc Bui, Soyeon Caren Han, Josiah Poon
Scene Graph Generation (SGG) serves a comprehensive representation of the images for human understanding as well as visual understanding tasks. Due to the long tail bias problem of…
SUPER-Rec: SUrrounding Position-Enhanced Representation for Recommendation
Taejun Lim, Siqu Long, Josiah Poon +1
Collaborative filtering problems are commonly solved based on matrix completion techniques which recover the missing values of user-item interaction matrices. In a matrix, the rati…
K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment
Jean Lee, Taejun Lim, Heejun Lee +4
Online hate speech detection has become an important issue due to the growth of online content, but resources in languages other than English are extremely limited. We introduce K-…
Understanding Attention for Vision-and-Language Tasks
Feiqi Cao, Soyeon Caren Han, Siqu Long +2
Attention mechanism has been used as an important component across Vision-and-Language(VL) tasks in order to bridge the semantic gap between visual and textual features. While atte…
Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis
Siwen Luo, Yihao Ding, Siqu Long +2
Recognizing the layout of unstructured digital documents is crucial when parsing the documents into the structured, machine-readable format for downstream applications. Recent stud…