1 citations · 1 across the 9 of their papers we have counts for
3 papers · 1 filter
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
Ishant Chintapatla, Kazuma Choji, Naaisha Agarwal +6
Recently, many benchmarks and datasets have been developed to evaluate Vision-Language Models (VLMs) using visual question answering (VQA) pairs, and models have shown significant…
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
Daniel Csizmadia, Andrei Codreanu, Victor Sim +5
We present Distill CLIP (DCLIP), a fine-tuned variant of the CLIP model that enhances multimodal image-text retrieval while preserving the original model's strong zero-shot classif…
Deconstructing Bias: A Multifaceted Framework for Diagnosing Cultural and Compositional Inequities in Text-to-Image Generative Models
Muna Numan Said, Aarib Zaidi, Rabia Usman +5
The transformative potential of text-to-image (T2I) models hinges on their ability to synthesize culturally diverse, photorealistic images from textual prompts. However, these mode…