collaborators

5 papers

cs.CV2025

Towards Efficient Exemplar Based Image Editing with Multimodal VLMs

Avadhoot Jadhav, Ashutosh Srivastava, Abhinav Java +4

Text-to-Image Diffusion models have enabled a wide array of image editing applications. However, capturing all types of edits through text alone can be challenging and cumbersome.…

cs.CV2025

HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs

Nikitha SR, Aradhya Neeraj Mathur, Tarun Ram Menta +2

The integration of high-resolution image features in modern multimodal large language models has demonstrated significant improvements in fine-grained visual understanding tasks, a…

cs.CV2025

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment

Nikitha SR, Tarun Ram Menta, Mausoom Sarkar

The advent of multimodal learning has brought a significant improvement in document AI. Documents are now treated as multimodal entities, incorporating both textual and visual info…

cs.LG2025

Analyzing Memorization in Large Language Models through the Lens of Model Attribution

Tarun Ram Menta, Susmit Agrawal, Chirag Agarwal

Large Language Models (LLMs) are prevalent in modern applications but often memorize training data, leading to privacy breaches and copyright issues. Existing research has mainly f…

cs.CV2024

ReEdit: Multimodal Exemplar-Based Image Editing with Diffusion Models

Ashutosh Srivastava, Tarun Ram Menta, Abhinav Java +4

Modern Text-to-Image (T2I) Diffusion models have revolutionized image editing by enabling the generation of high-quality photorealistic images. While the de facto method for perfor…