activity
20242026
collaborators

6 papers

cs.CV2026

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models

Nitish Shukla, Surgan Jandial, Arun Ross

Vision-Language Models (VLMs) have demonstrated remarkable progress in single-image understanding, yet effective reasoning across multiple images remains challenging. We identify a…

cs.CV2025

Towards Efficient Exemplar Based Image Editing with Multimodal VLMs

Avadhoot Jadhav, Ashutosh Srivastava, Abhinav Java +4

Text-to-Image Diffusion models have enabled a wide array of image editing applications. However, capturing all types of edits through text alone can be challenging and cumbersome.…

cs.CV2025

LEAST: "Local" text-conditioned image style transfer

Silky Singh, Surgan Jandial, Simra Shahid +1

Text-conditioned style transfer enables users to communicate their desired artistic styles through text descriptions, offering a new and expressive means of achieving stylization.…

cs.CV2024

ReEdit: Multimodal Exemplar-Based Image Editing with Diffusion Models

Ashutosh Srivastava, Tarun Ram Menta, Abhinav Java +4

Modern Text-to-Image (T2I) Diffusion models have revolutionized image editing by enabling the generation of high-quality photorealistic images. While the de facto method for perfor…

cs.CL2024

Thinking Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models

Shaz Furniturewala, Surgan Jandial, Abhinav Java +4

Existing debiasing techniques are typically training-based or require access to the model's internals and output distributions, so they are inaccessible to end-users looking to ada…

cs.CL2024

Introducing v0.5 of the AI Safety Benchmark from MLCommons

Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +97

This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safe…