activity
20182026
most citedFoundational Models Defining a New Era in Vision: A Survey and Outlook

68 citations · 130 across the 69 of their papers we have counts for

collaborators

85 papers

cs.CV2026

SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

Pengfei Li, Naufal Suryanto, Sicheng Zhang +2

Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understa…

cs.CV2026

CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation

Adonay Demewez Gebremedhin, Wessam Shehieb, Sara Alansari +4

Recent radiology multi-modal language models have made substantial progress in chest X-ray report generation, visual question answering, and temporal reasoning. While longitudinal…

cs.CV2026

From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

Basit Alawode, Moshira Ali Abdalla, Dwarikanath Mahapatra +2

Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution W…

cs.CL2026

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation

Sicheng Zhang, Zhonghao Yan, Binzhu Xie +4

Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual perf…

cs.CV2026

ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models

Hashmat Shadab Malik, Toluwani Aremu, Samuele Poppi +2

Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the pretrained model. However,…

cs.CV2026

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed +3

Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only by its final answer…