13 papers
A Test-time Actor-Critic Approach to News Images Generation
Damianos Galanopoulos, Vasileios Mezaris
This paper introduces the CERTH-ITI solution for the MediaEval NewsImages 2026 challenge, which focuses on generating images related to news headlines. Inspired by the Actor-Critic…
LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models
Nikolaos Gkalelis, Vasileios Mezaris
Large Vision-Language Models (VLMs) are successful in addressing a multitude of vision-language understanding tasks, such as Visual Question Answering (VQA), but their memory and c…
Sens-VisualNews: A Benchmark Dataset for Sensational Image Detection
Andreas Goulas, Damianos Galanopoulos, Evlampios Apostolidis +1
The detection of sensational content in media items can be a critical filtering mechanism for identifying check-worthy content and flagging potential disinformation, since such con…
SD-MVSum: Script-Driven Multimodal Video Summarization Method and Datasets
Manolis Mylonas, Charalampia Zerva, Evlampios Apostolidis +1
In this work, we present a method and two large-scale datasets for Script-Driven Multimodal Video Summarization. The proposed method, SD-MVSum, builds on our earlier SD-VSum method…
SDAKD: Student Discriminator Assisted Knowledge Distillation for Super-Resolution Generative Adversarial Networks
Nikolaos Kaparinos, Vasileios Mezaris
Generative Adversarial Networks (GANs) achieve excellent performance in generative tasks, such as image super-resolution, but their computational requirements make difficult their…
An Experimental Study on Generating Plausible Textual Explanations for Video Summarization
Thomas Eleftheriadis, Evlampios Apostolidis, Vasileios Mezaris
In this paper, we present our experimental study on generating plausible textual explanations for the outcomes of video summarization. For the needs of this study, we extend an exi…