4 citations · 4 across the 1 of their papers we have counts for
2 papers
cs.CV2023
VideoAdviser: Video Knowledge Distillation for Multimodal Transfer Learning
Yanan Wang, Donghuo Zeng, Shinya Wada +1
Multimodal transfer learning aims to transform pretrained representations of diverse modalities into a common domain space for effective multimodal fusion. However, conventional sy…
cs.CV2022★ 4 cited
VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question Answering
Yanan Wang, Michihiro Yasunaga, Hongyu Ren +2
Visual question answering (VQA) requires systems to perform concept-level reasoning by unifying unstructured (e.g., the context in question and answer; "QA context") and structured…