1 citations · 1 across the 7 of their papers we have counts for
5 papers · 1 filter
RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
Xing Zi, Jinghao Xiao, Yunxiao Shi +4
Visual Question Answering (VQA) in remote sensing (RS) is pivotal for interpreting Earth observation data. However, existing RS VQA datasets are constrained by limitations in annot…
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
Ali Anaissi, Junaid Akram, Kunal Chaturvedi +1
Memes are widely used for humor and cultural commentary, but they are increasingly exploited to spread hateful content. Due to their multimodal nature, hateful memes often evade tr…
DualPrompt-MedCap: A Dual-Prompt Enhanced Approach for Medical Image Captioning
Yining Zhao, Ali Braytee, Mukesh Prasad
Medical image captioning via vision-language models has shown promising potential for clinical diagnosis assistance. However, generating contextually relevant descriptions with acc…
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
Xing Zi, Tengjun Ni, Xianjing Fan +4
Accurate and automated captioning of aerial imagery is crucial for applications like environmental monitoring, urban planning, and disaster management. However, this task remains c…
Integration of Self-Supervised BYOL in Semi-Supervised Medical Image Recognition
Hao Feng, Yuanzhe Jia, Ruijia Xu +3
Image recognition techniques heavily rely on abundant labeled data, particularly in medical contexts. Addressing the challenges associated with obtaining labeled data has led to th…