379 citations · 451 across the 31 of their papers we have counts for
8 papers · 1 filter
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
David Romero, Chenyang Lyu, Haryo Akbarianto Wibowo +73
Visual Question Answering (VQA) is an important task in multimodal AI, and it is often used to test the ability of vision-language models to understand and reason on knowledge pres…
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
Oana Ignat, Longju Bai, Joan Nwatu +1
Current foundation models have shown impressive performance across various tasks. However, several studies have revealed that these models are not effective for everyone due to the…
Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
Oana Ignat, Santiago Castro, Weiji Li +1
We introduce the task of automatic human action co-occurrence identification, i.e., determine whether two human actions can co-occur in the same interval of time. We create and mak…
Scalable Performance Analysis for Vision-Language Models
Santiago Castro, Oana Ignat, Rada Mihalcea
Joint vision-language models have shown great performance over a diverse set of tasks. However, little is known about their limitations, as the high dimensional space learned by th…
WildQA: In-the-Wild Video Question Answering
Santiago Castro, Naihao Deng, Pingxuan Huang +2
Existing video understanding datasets mostly focus on human interactions, with little attention being paid to the "in the wild" settings, where the videos are recorded outdoors. We…
When Did It Happen? Duration-informed Temporal Localization of Narrated Actions in Vlogs
Oana Ignat, Santiago Castro, Yuhang Zhou +3
We consider the task of temporal human action localization in lifestyle vlogs. We introduce a novel dataset consisting of manual annotations of temporal localization for 13,000 nar…