collaborators

5 papers

cs.CV2026

Understanding Task Transfer in Vision-Language Models

Bhuvan Sachdeva, Karan Uppal, Abhinav Java +1

Vision-Language Models (VLMs) perform well on multimodal benchmarks but lag behind humans and specialized models on visual perception tasks like depth estimation or object counting…

cs.CL2026

FrugalRAG: Less is More in RL Finetuning for Multi-Hop Question Answering

Abhinav Java, Srivathsan Koundinyan, Nagarajan Natarajan +1

Reinforcement learning (RL) based on the final answer's reward has driven recent progress in small language models (SLMs) on reasoning-heavy tasks such as math and code. However, a…

cs.CL2025

Characterizing Deep Research: A Benchmark and Formal Definition

Abhinav Java, Ashmit Khandelwal, Sukruta Midigeshi +6

Information tasks such as writing surveys or analytical reports require complex search and reasoning, and have recently been grouped under the umbrella of \textit{deep research} --…

cs.CV2025

Towards Efficient Exemplar Based Image Editing with Multimodal VLMs

Avadhoot Jadhav, Ashutosh Srivastava, Abhinav Java +4

Text-to-Image Diffusion models have enabled a wide array of image editing applications. However, capturing all types of edits through text alone can be challenging and cumbersome.…

cs.CV2025

LEAST: "Local" text-conditioned image style transfer

Silky Singh, Surgan Jandial, Simra Shahid +1

Text-conditioned style transfer enables users to communicate their desired artistic styles through text descriptions, offering a new and expressive means of achieving stylization.…