12 papers
CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed +3
Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only by its final answer…
Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards
Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar +5
Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations,…
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar +4
Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and se…
DocAtlas: Multilingual Document Understanding Across 80+ Languages
Ahmed Heakl, Youssef Mohamed, Abdullah Sohail +6
Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that perpetuate existing biases. We…
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
Abdelrahman Shaker, Muhammad Maaz, Chenhui Gou +3
Video understanding models often struggle with high computational requirements, extensive parameter counts, and slow inference speed, making them inefficient for practical use. To…
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
Muhammad Maaz, Hanoona Rasheed, Fahad Shahbaz Khan +1
Reasoning over dynamic visual content remains a central challenge for multimodal large language models. Recent thinking models generate explicit reasoning traces for interpretabili…