9 papers
Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots
Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar +2
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alt…
Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards
Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar +5
Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations,…
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar +4
Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and se…
CoVR-R:Reason-Aware Composed Video Retrieval
Omkar Thawakar, Dmitry Demidov, Vaishnav Potlapalli +5
Composed Video Retrieval (CoVR) aims to find a target video given a reference video and a textual modification. Prior work assumes the modification text fully specifies the visual…
AgriChain Visually Grounded Expert Verified Reasoning for Interpretable Agricultural Vision Language Models
Hazza Mahmood, Yongqiang Yu, Rao Anwer
Accurate and interpretable plant disease diagnosis remains a major challenge for vision-language models (VLMs) in real-world agriculture. We introduce AgriChain, a dataset of appro…
MediX-R1: Open Ended Medical Reinforcement Learning
Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Omair Mohamed +5
We introduce MediX-R1, an open-ended Reinforcement Learning (RL) framework for medical multimodal large language models (MLLMs) that enables clinically grounded, free-form answers…