2 papers
cs.CL2026
MUNIChus: Multilingual News Image Captioning Benchmark
Yuji Chen, Alistair Plum, Hansi Hettiarachchi +4
The goal of news image captioning is to generate captions by integrating news article content with corresponding images, highlighting the relationship between textual context and v…
cs.LG2025
Evaluating Open-Source Vision-Language Models for Multimodal Sarcasm Detection
Saroj Basnet, Shafkat Farabi, Tharindu Ranasinghe +2
Recent advances in open-source vision-language models (VLMs) offer new opportunities for understanding complex and subjective multimodal phenomena such as sarcasm. In this work, we…