4 papers · 1 filter
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
Samarth Mishra, Kate Saenko, Venkatesh Saligrama
Compositionality, or correctly recognizing scenes as compositions of atomic visual concepts, remains difficult for multimodal large language models (MLLMs). Even state of the art M…
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
Kevin Miller, Samarth Mishra, Aditya Gangrade +2
Zero-shot multi-label recognition (MLR) with Vision-Language Models (VLMs) faces significant challenges without training data, model tuning, or architectural modifications. Existin…
Concept Arithmetics for Circumventing Concept Inhibition in Diffusion Models
Vitali Petsiuk, Kate Saenko
Motivated by ethical and legal concerns, the scientific community is actively developing methods to limit the misuse of Text-to-Image diffusion models for reproducing copyrighted,…
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
Samarth Mishra, Carlos D. Castillo, Hongcheng Wang +2
In cross-domain retrieval, a model is required to identify images from the same semantic category across two visual domains. For instance, given a sketch of an object, a model need…