7 papers
Causal Physics Steering in Video World Models via Concept Activation Vectors
Nahid Alam
Video world models learn representations of physical dynamics, but controlling their physical expectations at inference time remains an open problem. Recent interpretability work i…
Spatial Reasoning is Not a Free Lunch: A Controlled Study on LLaVA
Nahid Alam, Leema Krishna Murali, Siddhant Bharadwaj +7
Vision-language models (VLMs) have advanced rapidly, yet they still struggle with basic spatial reasoning. Despite strong performance on general benchmarks, modern VLMs remain brit…
Seeing Straight: Document Orientation Detection for Efficient OCR
Suranjan Goswami, Abhinav Ravi, Raja Kolla +5
Despite significant advances in document understanding, determining the correct orientation of scanned or photographed documents remains a critical pre-processing step in the real…
The Spatial Blindspot of Vision-Language Models
Nahid Alam, Leema Krishna Murali, Siddhant Bharadwaj +7
Vision-language models (VLMs) have advanced rapidly, but their ability to capture spatial relationships remains a blindspot. Current VLMs are typically built with contrastive langu…
Behind Maya: Building a Multilingual Vision Language Model
Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda +16
In recent times, we have seen a rapid development of large Vision-Language Models (VLMs). They have shown impressive results on academic benchmarks, primarily in widely spoken lang…
Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
Karthik Reddy Kanjula, Surya Guthikonda, Nahid Alam +1
Pretraining datasets are foundational to the development of multimodal models, yet they often have inherent biases and toxic content from the web-scale corpora they are sourced fro…