activity
20212024
most citedMeta Learning to Bridge Vision and Language Models for Multimodal Few-Shot Learning

8 citations · 18 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2023

Causalainer: Causal Explainer for Automatic Video Summarization

Jia-Hong Huang, Chao-Han Huck Yang, Pin-Yu Chen +2

The goal of video summarization is to automatically shorten videos such that it conveys the overall story without losing relevant information. In many application scenarios, improp…

cs.CV20231 cited

Improving Visual Question Answering Models through Robustness Analysis and In-Context Learning with a Chain of Basic Questions

Jia-Hong Huang, Modar Alfadly, Bernard Ghanem +1

Deep neural networks have been critical in the task of Visual Question Answering (VQA), with research traditionally focused on improving model accuracy. Recently, however, there ha…

cs.CV20238 cited

Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot Learning

Ivona Najdenkoska, Xiantong Zhen, Marcel Worring

Multimodal few-shot learning is challenging due to the large domain gap between vision and language modalities. Existing methods are trying to communicate visual concepts as prompt…

cs.CV20237 cited

X-TRA: Improving Chest X-ray Tasks with Cross-Modal Retrieval Augmentation

Tom van Sonsbeek, Marcel Worring

An important component of human analysis of medical images and their context is the ability to relate newly seen things to related instances in our memory. In this paper we mimic t…

cs.CV2022

PanorAMS: Automatic Annotation for Detecting Objects in Urban Context

Inske Groenen, Stevan Rudinac, Marcel Worring

Large collections of geo-referenced panoramic images are freely available for cities across the globe, as well as detailed maps with location and meta-data on a great variety of ur…