7 papers
LiteEmbed: Adapting CLIP to Rare Classes
Aishwarya Agarwal, Srikrishna Karanam, Vineet Gandhi
Large-scale vision-language models such as CLIP achieve strong zero-shot recognition but struggle with classes that are rarely seen during pretraining, including newly emerging ent…
Learning 3D Texture-Aware Representations for Parsing Diverse Human Clothing and Body Parts
Kiran Chhatre, Christopher Peters, Srikrishna Karanam
Existing methods for human parsing into body parts and clothing often use fixed mask categories with broad labels that obscure fine-grained clothing types. Recent open-vocabulary s…
Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
Aishwarya Agarwal, Srikrishna Karanam, Vineet Gandhi
Contrastive vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition yet remain vulnerable to spurious correlations, particularly background over-reliance. W…
Composing Parts for Expressive Object Generation
Harsh Rangwani, Aishwarya Agarwal, Kuldeep Kulkarni +2
Image composition and generation are processes where the artists need control over various parts of the generated images. However, the current state-of-the-art generation models, l…
TIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time Correction
Aishwarya Agarwal, Srikrishna Karanam, Vineet Gandhi
We consider the problem of single-source domain generalization. Existing methods typically rely on extensive augmentations to synthetically cover diverse domains during training. H…
CoCoNO: Attention Contrast-and-Complete for Initial Noise Optimization in Text-to-Image Synthesis
Aravindan Sundaram, Ujjayan Pal, Abhimanyu Chauhan +2
Despite recent advancements in text-to-image models, achieving semantically accurate images in text-to-image diffusion models is a persistent challenge. While existing initial late…