Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning
Chashi Mahiul Islam, Oteo Mamo, Samuel Jacob Chacko +2
Vision-language models (VLMs) have advanced multimodal reasoning but still face challenges in spatial reasoning for 3D scenes and complex object configurations. To address this, we…
cs.CV2025
DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities
Chashi Mahiul Islam, Samuel Jacob Chacko, Preston Horne +1
Multimodal Large Language Models (MLLMs) represent the cutting edge of AI technology, with DeepSeek models emerging as a leading open-source alternative offering competitive perfor…
cs.CV2025
Mechanistic Understandings of Representation Vulnerabilities and Engineering Robust Vision Transformers
Chashi Mahiul Islam, Samuel Jacob Chacko, Mao Nishino +1
While transformer-based models dominate NLP and vision applications, their underlying mechanisms to map the input space to the label space semantically are not well understood. In…