32 papers
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
Sicheng Zhang, Zhonghao Yan, Binzhu Xie +4
Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual perf…
From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
Basit Alawode, Moshira Ali Abdalla, Dwarikanath Mahapatra +3
Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution W…
ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models
Hashmat Shadab Malik, Toluwani Aremu, Samuele Poppi +2
Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However,…
CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed +3
Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only by its final answer…
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
Sicheng Zhang, Muzammal Naseer, Binzhu Xie +5
CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applicat…
SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking
Mohamad Alansari, Yonathan Michael, Hasan AlMarzouqi +3
We revisit the memory update mechanism in SAM2-based visual object tracking and identify confidence-only mask selection as the dominant cause of drift under occlusion, rapid motion…