8 citations · 18 across the 10 of their papers we have counts for
5 papers · 1 filter
Causalainer: Causal Explainer for Automatic Video Summarization
Jia-Hong Huang, Chao-Han Huck Yang, Pin-Yu Chen +2
The goal of video summarization is to automatically shorten videos such that it conveys the overall story without losing relevant information. In many application scenarios, improp…
Improving Visual Question Answering Models through Robustness Analysis and In-Context Learning with a Chain of Basic Questions
Jia-Hong Huang, Modar Alfadly, Bernard Ghanem +1
Deep neural networks have been critical in the task of Visual Question Answering (VQA), with research traditionally focused on improving model accuracy. Recently, however, there ha…
Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot Learning
Ivona Najdenkoska, Xiantong Zhen, Marcel Worring
Multimodal few-shot learning is challenging due to the large domain gap between vision and language modalities. Existing methods are trying to communicate visual concepts as prompt…
X-TRA: Improving Chest X-ray Tasks with Cross-Modal Retrieval Augmentation
Tom van Sonsbeek, Marcel Worring
An important component of human analysis of medical images and their context is the ability to relate newly seen things to related instances in our memory. In this paper we mimic t…
PanorAMS: Automatic Annotation for Detecting Objects in Urban Context
Inske Groenen, Stevan Rudinac, Marcel Worring
Large collections of geo-referenced panoramic images are freely available for cities across the globe, as well as detailed maps with location and meta-data on a great variety of ur…