9 citations · 9 across the 3 of their papers we have counts for
4 papers · 1 filter
SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
Samarth Mishra, Kate Saenko, Venkatesh Saligrama
Compositionality, or correctly recognizing scenes as compositions of atomic visual concepts, remains difficult for multimodal large language models (MLLMs). Even state of the art M…
SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models
Kevin Miller, Samarth Mishra, Aditya Gangrade +2
Zero-shot multi-label recognition (MLR) with Vision-Language Models (VLMs) faces significant challenges without training data, model tuning, or architectural modifications. Existin…
Learning Compositional Representations for Effective Low-Shot Generalization
Samarth Mishra, Pengkai Zhu, Venkatesh Saligrama
We propose Recognition as Part Composition (RPC), an image encoding approach inspired by human cognition. It is based on the cognitive theory that humans recognize complex objects…
Effectively Leveraging Attributes for Visual Similarity
Samarth Mishra, Zhongping Zhang, Yuan Shen +3
Measuring similarity between two images often requires performing complex reasoning along different axes (e.g., color, texture, or shape). Insights into what might be important for…