collaborators

32 papers

cs.CL2026

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation

Sicheng Zhang, Zhonghao Yan, Binzhu Xie +4

Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual perf…

cs.CV2026

From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

Basit Alawode, Moshira Ali Abdalla, Dwarikanath Mahapatra +3

Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution W…

cs.CV2026

ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models

Hashmat Shadab Malik, Toluwani Aremu, Samuele Poppi +2

Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However,…

cs.CV2026

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed +3

Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only by its final answer…

cs.CV2026

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

Sicheng Zhang, Muzammal Naseer, Binzhu Xie +5

CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applicat…

cs.CV2026

SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking

Mohamad Alansari, Yonathan Michael, Hasan AlMarzouqi +3

We revisit the memory update mechanism in SAM2-based visual object tracking and identify confidence-only mask selection as the dominant cause of drift under occlusion, rapid motion…