collaborators

15 papers

cs.CV2026

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots

Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar +2

Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alt…

cs.CV2026

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar +5

Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations,…

cs.CV2026

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

Omkar Thawakar, Shravan Venkatraman, Ritesh Thawkar +5

Recent advances in large multimodal models (LMMs) have enabled impressive reasoning and perception abilities, yet most existing training pipelines still depend on human-curated dat…

cs.CV2026

Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos

Bonan Ding, Umair Nawaz, Ufaq Khan +5

Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evi…

cs.LG2026

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

Ahmed Heakl, Abdelrahman M. Shaker, Youssef Mohamed +4

When a model produces a correct solution under reinforcement learning with verifiable rewards (RLVR), every token receives the same reward signal regardless of whether it was a dec…

cs.CV2026

WorldCache: Content-Aware Caching for Accelerated Video World Models

Umair Nawaz, Ahmed Heakl, Ufaq Khan +3

Diffusion Transformers (DiTs) power high-fidelity video world models but remain computationally expensive due to sequential denoising and costly spatio-temporal attention. Training…