activity
20172026
most citedTowards Robust RGB-D Human Mesh Recovery

11 citations · 25 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

24 papers · 1 filter

cs.CV2026

RefDiT: Local Attribute Guidance in Reference-Based Image Generation

Rameshwar Mishra, Srikrishna Karanam, A V Subramanyam

Personalization models generate new images guided by a few subject references, while style transfer methods aim to produce images aligned with a global style derived from a referen…

cs.CV2024

SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation

Sayan Nag, Koustava Goswami, Srikrishna Karanam

Referring Expression Segmentation (RES) aims to provide a segmentation mask of the target object in an image referred to by the text (i.e., referring expression). Existing methods…

cs.CV2024

AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models

Aishwarya Agarwal, Srikrishna Karanam, Balaji Vasan Srinivasan

We consider the problem of customizing text-to-image diffusion models with user-supplied reference images. Given new prompts, the existing methods can capture the key concept from…

cs.CV2024

Composing Parts for Expressive Object Generation

Harsh Rangwani, Aishwarya Agarwal, Kuldeep Kulkarni +2

Image composition and generation are processes where the artists need control over various parts of the generated images. However, the current state-of-the-art generation models, l…

cs.CV2024

Few Shot Class Incremental Learning using Vision-Language models

Anurag Kumar, Chinmay Bharti, Saikat Dutta +2

Recent advancements in deep learning have demonstrated remarkable performance comparable to human capabilities across various supervised computer vision tasks. However, the prevale…

cs.CV2023

Approximate Caching for Efficiently Serving Diffusion Models

Shubham Agarwal, Subrata Mitra, Sarthak Chakraborty +3

Text-to-image generation using diffusion models has seen explosive popularity owing to their ability in producing high quality images adhering to text prompts. However, production-…