8 citations · 10 across the 9 of their papers we have counts for
3 papers · 1 filter
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
Shuo Yang, Siwen Luo, Soyeon Caren Han +1
Visual Question Answering (VQA) requires reasoning across visual and textual modalities, yet Large Vision-Language Models (LVLMs) often lack integrated commonsense knowledge, limit…
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
Shuo Yang, Siwen Luo, Soyeon Caren Han
Existing Multimodal Large Language Models (MLLMs) and Visual Language Pretrained Models (VLPMs) have shown remarkable performances in the general Visual Question Answering (VQA). H…
3M-Health: Multimodal Multi-Teacher Knowledge Distillation for Mental Health Detection
Rina Carines Cabral, Siwen Luo, Josiah Poon +1
The significance of mental health classification is paramount in contemporary society, where digital platforms serve as crucial sources for monitoring individuals' well-being. Howe…