collaborators

9 papers

cs.CV2026

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots

Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar +2

Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alt…

cs.CV2026

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar +5

Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations,…

cs.CV2026

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar +4

Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and se…

cs.CV2026

CoVR-R:Reason-Aware Composed Video Retrieval

Omkar Thawakar, Dmitry Demidov, Vaishnav Potlapalli +5

Composed Video Retrieval (CoVR) aims to find a target video given a reference video and a textual modification. Prior work assumes the modification text fully specifies the visual…

cs.CV2026

AgriChain Visually Grounded Expert Verified Reasoning for Interpretable Agricultural Vision Language Models

Hazza Mahmood, Yongqiang Yu, Rao Anwer

Accurate and interpretable plant disease diagnosis remains a major challenge for vision-language models (VLMs) in real-world agriculture. We introduce AgriChain, a dataset of appro…

cs.CV2026

MediX-R1: Open Ended Medical Reinforcement Learning

Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Omair Mohamed +5

We introduce MediX-R1, an open-ended Reinforcement Learning (RL) framework for medical multimodal large language models (MLLMs) that enables clinically grounded, free-form answers…