works on

From the 1 of 16 linked papers with an AI index.

activity
20242026
most citedDistractor-Aware Video Object Segmentation

3 citations · 3 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV20263 cited

Distractor-Aware Video Object Segmentation

Andreas Robinson, Abdelrahman Eldesokey, Michael Felsberg

Semi-supervised video object segmentation is a challenging task that aims to segment a target throughout a video sequence given an initial mask at the first frame. Discriminative a…

cs.CV2026

DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation

Qian Wang, Zhenyu Li, Abdelrahman Eldesokey +1

Subject-driven image generation faces an "Identity-Diversity Paradox", where strong identity preservation often leads to rigid and low-diversity outputs. We propose a post-training…

cs.CV2026

CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models

Reem Alzahrani, Hassan Alshanqiti, Bushra Bin Hemid +3

Vision-Language Models (VLMs) excel at multimodal reasoning, yet it remains unclear whether their answers are grounded in visual evidence or driven by learned language and world pr…

cs.CV2026

Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation

Abdelrahman Eldesokey, Merey Ramazanova, Ahmad Sait +4

Text-to-image (T2I) generation has advanced rapidly, making reliable evaluation critical as performance differences between models narrow. Existing evaluation practices typically a…

cs.CV2026

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

Ismael Elsharkawi, Ahmed Sait, Silvio Giancola +3

Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos due to large viewpoint variati…

cs.CV2026

NearID: Identity Representation Learning via Near-identity Distractors

Aleksandar Cvejic, Rameen Abdal, Abdelrahman Eldesokey +2

When evaluating identity-focused tasks such as personalized generation and image editing, existing vision encoders entangle object identity with background context, leading to unre…