11 papers
Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards
Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar +5
Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations,…
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar +4
Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and se…
PCM-NeRF: Probabilistic Camera Modeling for Neural Radiance Fields under Pose Uncertainty
Shravan Venkatraman, Rakesh Raj Madavan, Pavan Kumar Sathya Venkatesh
Neural surface reconstruction methods typically treat camera poses as fixed values, assuming perfect accuracy from Structure-from-Motion (SfM) systems. This assumption breaks down…
TIDE: Two-Stage Inverse Degradation Estimation with Guided Prior Disentanglement for Underwater Image Restoration
Shravan Venkatraman, Rakesh Raj Madavan, Pavan Kumar S +1
Underwater image restoration is essential for marine applications ranging from ecological monitoring to archaeological surveys, but effectively addressing the complex and spatially…
UGPL: Uncertainty-Guided Progressive Learning for Evidence-Based Classification in Computed Tomography
Shravan Venkatraman, Pavan Kumar S, Rakesh Raj Madavan +1
Accurate classification of computed tomography (CT) images is essential for diagnosis and treatment planning, but existing methods often struggle with the subtle and spatially dive…
FUSION: Frequency-guided Underwater Spatial Image recOnstructioN
Jaskaran Singh Walia, Shravan Venkatraman, Pavithra LK
Underwater images suffer from severe degradations, including color distortions, reduced visibility, and loss of structural details due to wavelength-dependent attenuation and scatt…