activity
20212026
most citedChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

40 citations · 43 across the 5 of their papers we have counts for

collaborators

7 papers

cs.LG2026

A Shared Valence Axis Across Modern LLMs and Human EEG: The Saturation Regularity

Yousef A. Radwan, Xuhui Liu, Kilichbek Haydarov +2

Large language models (LLMs) have emerged as powerful representation learners whose internal features increasingly align with human cognition. We study whether modern LLMs can serv…

cs.CL2024★ 1 cited

No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages

Youssef Mohamed, Runjia Li, Ibrahim Said Ahmad +4

Research in vision and language has made considerable progress thanks to benchmarks such as COCO. COCO captions focused on unambiguous facts in English; ArtEmis introduced subjecti…

cs.CL2023

Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations

Kilichbek Haydarov, Xiaoqian Shen, Avinash Madasu +4

We introduce Affective Visual Dialog, an emotion explanation and reasoning task as a testbed for research on understanding the formation of emotions in visually grounded conversati…

cs.CV2023

Video ChatCaptioner: Towards Enriched Spatiotemporal Descriptions

Jun Chen, Deyao Zhu, Kilichbek Haydarov +2

Video captioning aims to convey dynamic scenes from videos using natural language, facilitating the understanding of spatiotemporal information within our environment. Although the…

cs.CV2023★ 40 cited

ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Deyao Zhu, Jun Chen, Kilichbek Haydarov +3

Asking insightful questions is crucial for acquiring knowledge and expanding our understanding of the world. However, the importance of questioning has been largely overlooked in A…

cs.CV2022★ 2 cited

It is Okay to Not Be Okay: Overcoming Emotional Bias in Affective Image Captioning by Contrastive Data Collection

Youssef Mohamed, Faizan Farooq Khan, Kilichbek Haydarov +1

Datasets that capture the connection between vision, language, and affection are limited, causing a lack of understanding of the emotional aspect of human intelligence. As a step i…