collaborators

12 papers

cs.CV2026

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed +3

Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only by its final answer…

cs.CV2026

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar +5

Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations,…

cs.CV2026

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar +4

Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and se…

cs.CL2026

DocAtlas: Multilingual Document Understanding Across 80+ Languages

Ahmed Heakl, Youssef Mohamed, Abdullah Sohail +6

Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that perpetuate existing biases. We…

cs.CV2026

Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding

Abdelrahman Shaker, Muhammad Maaz, Chenhui Gou +3

Video understanding models often struggle with high computational requirements, extensive parameter counts, and slow inference speed, making them inefficient for practical use. To…

cs.CV2025

Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models

Muhammad Maaz, Hanoona Rasheed, Fahad Shahbaz Khan +1

Reasoning over dynamic visual content remains a central challenge for multimodal large language models. Recent thinking models generate explicit reasoning traces for interpretabili…