activity
20242026
most citedBeyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language

1 citations · 2 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

Tell me Habibi, is it Real or Fake?

Kartik Kuckreja, Parul Gupta, Injy Hamed +3

Deepfake generation methods are evolving fast, making fake media harder to detect and raising serious societal concerns. Most deepfake detection and dataset creation research focus…

cs.CV2025

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana +66

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cul…

cs.CV2024

CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

David Romero, Chenyang Lyu, Haryo Akbarianto Wibowo +73

Visual Question Answering (VQA) is an important task in multimodal AI, and it is often used to test the ability of vision-language models to understand and reason on knowledge pres…

cs.CV2024

Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering

David Romero, Thamar Solorio

We present Q-ViD, a simple approach for video question answering (video QA), that unlike prior methods, which are based on complex architectures, computationally expensive pipeline…

cs.CV2024

Labeling Comic Mischief Content in Online Videos with a Multimodal Hierarchical-Cross-Attention Model

Elaheh Baharlouei, Mahsa Shafaei, Yigeng Zhang +2

We address the challenge of detecting questionable content in online media, specifically the subcategory of comic mischief. This type of content combines elements such as violence,…