collaborators

6 papers

cs.CV2026

Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance

Pradhaan S Bhat, Rishubh Parihar, Abhijnya Bhat +1

State-of-the-art flow models generate stunning images from text or image prompts. However, they suffer from diversity collapse when generating multiple samples under the same condi…

cs.CV2026

Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping

Rishubh Parihar, Ayush Raina, R. Venkatesh Babu +1

Reference-based diffusion models enable highly controllable image generation by leveraging elements from input images to guide prompt-driven synthesis. However, these models are co…

cs.CV2026

Thinking in Boxes: 3D Editing in Real Images Made Easy

Pradhaan S Bhat, Naveen Chandra R, Rishubh Parihar +4

Text and 2D-conditioning interfaces provide weak, ambiguous control over spatial transformations in image editing -- particularly under large object motions and camera changes. Pri…

cs.CV2026

Adaptive Tokenisation Via Temporal Redundancy Masking And Latent Inpainting

Kevin Dave, Sai Aditya Patkuri, Chhaya Kumar Das +3

Adaptive video tokenisation seeks to dynamically allocate token budgets based on the underlying visual complexity of a sequence. Current continuous-regime approaches achieve this v…

cs.CV2026

Do Vision Language Models Need to Process Image Tokens?

Sambit Ghosh, R. Venkatesh Babu, Chirag Agarwal

Vision Language Models (VLMs) have achieved remarkable success by integrating visual encoders with large language models (LLMs). While VLMs process dense image tokens across deep t…

cs.CV2025

Composing Parts for Expressive Object Generation

Harsh Rangwani, Aishwarya Agarwal, Kuldeep Kulkarni +2

Image composition and generation are processes where the artists need control over various parts of the generated images. However, the current state-of-the-art generation models, l…