collaborators

11 papers

cs.CV2026

Attending to Multimodal Generation One Token at a Time

Varun Gupta, Vineet Gandhi, Makarand Tapaswi

Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic information in an evolving context. Prior work on interpretability h…

cs.CV2026

LiteEmbed: Adapting CLIP to Rare Classes

Aishwarya Agarwal, Srikrishna Karanam, Vineet Gandhi

Large-scale vision-language models such as CLIP achieve strong zero-shot recognition but struggle with classes that are rarely seen during pretraining, including newly emerging ent…

cs.CV2025

Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach

Aishwarya Agarwal, Srikrishna Karanam, Vineet Gandhi

Contrastive vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition yet remain vulnerable to spurious correlations, particularly background over-reliance. W…

cs.LG2025

Simplifying Knowledge Transfer in Pretrained Models

Siddharth Jain, Shyamgopal Karthik, Vineet Gandhi

Pretrained models are ubiquitous in the current deep learning landscape, offering strong results on a broad range of tasks. Recent works have shown that models differing in various…

cs.CV2025

Pseudo-labelling meets Label Smoothing for Noisy Partial Label Learning

Darshana Saravanan, Naresh Manwani, Vineet Gandhi

We motivate weakly supervised learning as an effective learning paradigm for problems where curating perfectly annotated datasets is expensive and may require domain expertise such…

cs.CV2025

Investigating Mechanisms for In-Context Vision Language Binding

Darshana Saravanan, Makarand Tapaswi, Vineet Gandhi

To understand a prompt, Vision-Language models (VLMs) must perceive the image, comprehend the text, and build associations within and across both modalities. For instance, given an…