papers

Publications (16)

cs.CL2019

Generating Natural Language Explanations for Visual Question Answering using Scene Graphs and Visual Attention

Shalini Ghosh, Giedrius Burachas, Arijit Ray +1

In this paper, we present a novel approach for the task of eXplainable Question Answering (XQA), i.e., generating natural language (NL) explanations for the Visual Question Answeri…

cs.CV2026

Mull-Tokens: Modality-Agnostic Latent Thinking

Arijit Ray, Ahmed Abdelkader, Chengzhi Mao +5

Reasoning goes beyond language; the real world requires reasoning about space, time, affordances, and much more that words alone cannot convey. Existing multimodal models exploring…

cs.RO2025

GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation

Abhay Deshpande, Yuquan Deng, Arijit Ray +7

We present GrasMolmo, a generalizable open-vocabulary task-oriented grasping (TOG) model. GraspMolmo predicts semantically appropriate, stable grasps conditioned on a natural langu…

cs.AI2023

Socratis: Are large multimodal models emotionally aware?

Katherine Deng, Arijit Ray, Reuben Tan +3

Existing emotion prediction benchmarks contain coarse emotion labels which do not consider the diversity of emotions that an image and text can elicit in humans due to various reas…

cs.CV2016

Question Relevance in VQA: Identifying Non-Visual And False-Premise Questions

Arijit Ray, Gordon Christie, Mohit Bansal +2

Visual Question Answering (VQA) is the task of answering natural-language questions about images. We introduce the novel problem of determining the relevance of questions to images…

cs.CV2021

Generating and Evaluating Explanations of Attended and Error-Inducing Input Regions for VQA Models

Arijit Ray, Michael Cogswell, Xiao Lin +4

Attention maps, a popular heatmap-based explanation method for Visual Question Answering (VQA), are supposed to help users understand the model by highlighting portions of the imag…

cs.CV2023

COLA: A Benchmark for Compositional Text-to-image Retrieval

Arijit Ray, Filip Radenovic, Abhimanyu Dubey +3

Compositional reasoning is a hallmark of human visual intelligence. Yet, despite the size of large vision-language models, they struggle to represent simple compositions by combini…

cs.CV2020

The Impact of Explanations on AI Competency Prediction in VQA

Kamran Alipour, Arijit Ray, Xiao Lin +3

Explainability is one of the key elements for building trust in AI systems. Among numerous attempts to make AI explainable, quantifying the effect of explanations remains a challen…

cs.CV2023

Language-Guided Audio-Visual Source Separation via Trimodal Consistency

Reuben Tan, Arijit Ray, Andrea Burns +5

We propose a self-supervised approach for learning to perform audio source separation in videos based on natural language queries, using only unlabeled video and audio pairs as tra…

cs.CV2021

Improving Users' Mental Model with Attention-directed Counterfactual Edits

Kamran Alipour, Arijit Ray, Xiao Lin +4

In the domain of Visual Question Answering (VQA), studies have shown improvement in users' mental model of the VQA system when they are exposed to examples of how these systems ans…

cs.CV2025

SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding

Ellis Brown, Arijit Ray, Ranjay Krishna +3

Despite impressive high-level video comprehension, multimodal language models struggle with spatial reasoning across time and space. While current spatial training approaches rely…

cs.CV2019

Sunny and Dark Outside?! Improving Answer Consistency in VQA through Entailed Question Generation

Arijit Ray, Karan Sikka, Ajay Divakaran +2

While models for Visual Question Answering (VQA) have steadily improved over the years, interacting with one quickly reveals that these models lack consistency. For instance, if a…

cs.CV2025

SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models

Arijit Ray, Jiafei Duan, Ellis Brown +9

Reasoning about motion and space is a fundamental cognitive capability that is required by multiple real-world applications. While many studies highlight that large multimodal lang…

cs.CV2024

BloomVQA: Assessing Hierarchical Multi-modal Comprehension

Yunye Gong, Robik Shrestha, Jared Claypoole +4

We propose a novel VQA dataset, BloomVQA, to facilitate comprehensive evaluation of large vision-language models on comprehension tasks. Unlike current benchmarks that often focus…

cs.CY2019

Can You Explain That? Lucid Explanations Help Human-AI Collaborative Image Retrieval

Arijit Ray, Yi Yao, Rakesh Kumar +2

While there have been many proposals on making AI algorithms explainable, few have attempted to evaluate the impact of AI-generated explanations on human performance in conducting…

cs.CV2023

Lasagna: Layered Score Distillation for Disentangled Object Relighting

Dina Bashkirova, Arijit Ray, Rupayan Mallick +4

Professional artists, photographers, and other visual content creators use object relighting to establish their photo's desired effect. Unfortunately, manual tools that allow relig…