3 papers
cs.CV2026
S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models
Nitish Shukla, Surgan Jandial, Arun Ross
Vision-Language Models (VLMs) have demonstrated remarkable progress in single-image understanding, yet effective reasoning across multiple images remains challenging. We identify a…
cs.CV2025
Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
Avadhoot Jadhav, Ashutosh Srivastava, Abhinav Java +4
Text-to-Image Diffusion models have enabled a wide array of image editing applications. However, capturing all types of edits through text alone can be challenging and cumbersome.…
cs.CV2025
LEAST: "Local" text-conditioned image style transfer
Silky Singh, Surgan Jandial, Simra Shahid +1
Text-conditioned style transfer enables users to communicate their desired artistic styles through text descriptions, offering a new and expressive means of achieving stylization.…