5 papers
The Spatial Blindspot of Vision-Language Models
Nahid Alam, Leema Krishna Murali, Siddhant Bharadwaj +7
Vision-language models (VLMs) have advanced rapidly, but their ability to capture spatial relationships remains a blindspot. Current VLMs are typically built with contrastive langu…
Behind Maya: Building a Multilingual Vision Language Model
Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda +16
In recent times, we have seen a rapid development of large Vision-Language Models (VLMs). They have shown impressive results on academic benchmarks, primarily in widely spoken lang…
Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
Karthik Reddy Kanjula, Surya Guthikonda, Nahid Alam +1
Pretraining datasets are foundational to the development of multimodal models, yet they often have inherent biases and toxic content from the web-scale corpora they are sourced fro…
Maya: An Instruction Finetuned Multilingual Multimodal Model
Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda +16
The rapid development of large Vision-Language Models (VLMs) has led to impressive results on academic benchmarks, primarily in widely spoken languages. However, significant gaps r…
Embedding Geometries of Contrastive Language-Image Pre-Training
Jason Chuan-Chih Chou, Nahid Alam
Since the publication of CLIP, the approach of using InfoNCE loss for contrastive pre-training has become widely popular for bridging two or more modalities. Despite its wide adopt…