From the 1 of 16 linked papers with an AI index.
3 citations · 3 across the 9 of their papers we have counts for
12 papers · 1 filter
Distractor-Aware Video Object Segmentation
Andreas Robinson, Abdelrahman Eldesokey, Michael Felsberg
Semi-supervised video object segmentation is a challenging task that aims to segment a target throughout a video sequence given an initial mask at the first frame. Discriminative a…
DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation
Qian Wang, Zhenyu Li, Abdelrahman Eldesokey +1
Subject-driven image generation faces an "Identity-Diversity Paradox", where strong identity preservation often leads to rigid and low-diversity outputs. We propose a post-training…
CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models
Reem Alzahrani, Hassan Alshanqiti, Bushra Bin Hemid +3
Vision-Language Models (VLMs) excel at multimodal reasoning, yet it remains unclear whether their answers are grounded in visual evidence or driven by learned language and world pr…
Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation
Abdelrahman Eldesokey, Merey Ramazanova, Ahmad Sait +4
Text-to-image (T2I) generation has advanced rapidly, making reliable evaluation critical as performance differences between models narrow. Existing evaluation practices typically a…
SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
Ismael Elsharkawi, Ahmed Sait, Silvio Giancola +3
Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos due to large viewpoint variati…
NearID: Identity Representation Learning via Near-identity Distractors
Aleksandar Cvejic, Rameen Abdal, Abdelrahman Eldesokey +2
When evaluating identity-focused tasks such as personalized generation and image editing, existing vision encoders entangle object identity with background context, leading to unre…