11 papers
ZIPP:Zero-shot Image Personalization from Personas
Harini SI, Somesh Singh, Yaman Kumar Singla +2
Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate aesthetics rather than indivi…
TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation
Pengyu Yan, Akhil Gorugantu, Mahesh Bhosale +3
Multi-video event understanding demands models that can locate and attribute query-relevant evidence scattered across long, heterogeneous video corpora. Existing large vision-langu…
Score-Control for Hallucination Reduction in Diffusion Models
Mahesh Bhosale, Naresh Kumar Devulapally, Abdul Wasi +3
Diffusion models have emerged as the backbone of modern generative AI, powering advances in vision, language, audio and other modalities. Despite their success, they suffer from ha…
CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering
Mahesh Bhosale, Abdul Wasi, Vishvesh Trivedi +3
Grounded multi-video question answering over real-world news events requires systems to surface query-relevant evidence across heterogeneous video archives while attributing every…
CWCD: Category-Wise Contrastive Decoding for Structured Medical Report Generation
Shantam Srivastava, Mahesh Bhosale, David Doermann +1
Interpreting chest X-rays is inherently challenging due to the overlap between anatomical structures and the subtle presentation of many clinically significant pathologies, making…
FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants
Mahesh Bhosale, Abdul Wasi, Shantam Srivastava +5
While powerful in image-conditioned generation, multimodal large language models (MLLMs) can display uneven performance across demographic groups, highlighting fairness risks. In s…