41 citations · 68 across the 10 of their papers we have counts for
12 papers
PiggyBack: Pretrained Visual Question Answering Environment for Backing up Non-deep Learning Professionals
Zhihao Zhang, Siwen Luo, Junyi Chen +4
We propose a PiggyBack, a Visual Question Answering platform that allows users to apply the state-of-the-art visual-language pretrained models easily. The PiggyBack supports the fu…
SUPER-Rec: SUrrounding Position-Enhanced Representation for Recommendation
Taejun Lim, Siqu Long, Josiah Poon +1
Collaborative filtering problems are commonly solved based on matrix completion techniques which recover the missing values of user-item interaction matrices. In a matrix, the rati…
Understanding Attention for Vision-and-Language Tasks
Feiqi Cao, Soyeon Caren Han, Siqu Long +2
Attention mechanism has been used as an important component across Vision-and-Language(VL) tasks in order to bridge the semantic gap between visual and textual features. While atte…
Doc-GCN: Heterogeneous Graph Convolutional Networks for Document Layout Analysis
Siwen Luo, Yihao Ding, Siqu Long +2
Recognizing the layout of unstructured digital documents is crucial when parsing the documents into the structured, machine-readable format for downstream applications. Recent stud…
Vision-and-Language Pretrained Models: A Survey
Siqu Long, Feiqi Cao, Soyeon Caren Han +1
Pretrained models have produced great success in both Computer Vision (CV) and Natural Language Processing (NLP). This progress leads to learning joint representations of vision an…
ME-GCN: Multi-dimensional Edge-Embedded Graph Convolutional Networks for Semi-supervised Text Classification
Kunze Wang, Soyeon Caren Han, Siqu Long +1
Compared to sequential learning models, graph-based neural networks exhibit excellent ability in capturing global information and have been used for semi-supervised learning tasks.…