activity
20182025
most citedA Study on Multimodal and Interactive Explanations for Visual Question Answering

8 citations · 10 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV20211 cited

Improving Users' Mental Model with Attention-directed Counterfactual Edits

Kamran Alipour, Arijit Ray, Xiao Lin +4

In the domain of Visual Question Answering (VQA), studies have shown improvement in users' mental model of the VQA system when they are exposed to examples of how these systems ans…

cs.CV20211 cited

Modular Adaptation for Cross-Domain Few-Shot Learning

Xiao Lin, Meng Ye, Yunye Gong +4

Adapting pre-trained representations has become the go-to recipe for learning new downstream tasks with limited examples. While literature has demonstrated great successes via repr…

cs.CV2021

Generating and Evaluating Explanations of Attended and Error-Inducing Input Regions for VQA Models

Arijit Ray, Michael Cogswell, Xiao Lin +4

Attention maps, a popular heatmap-based explanation method for Visual Question Answering (VQA), are supposed to help users understand the model by highlighting portions of the imag…

cs.CV2020

Hybrid Consistency Training with Prototype Adaptation for Few-Shot Learning

Meng Ye, Xiao Lin, Giedrius Burachas +2

Few-Shot Learning (FSL) aims to improve a model's generalization capability in low data regimes. Recent FSL works have made steady progress via metric learning, meta learning, repr…

cs.CV2020

The Impact of Explanations on AI Competency Prediction in VQA

Kamran Alipour, Arijit Ray, Xiao Lin +3

Explainability is one of the key elements for building trust in AI systems. Among numerous attempts to make AI explainable, quantifying the effect of explanations remains a challen…

cs.CV2018

Stacked Spatio-Temporal Graph Convolutional Networks for Action Segmentation

Pallabi Ghosh, Yi Yao, Larry S. Davis +1

We propose novel Stacked Spatio-Temporal Graph Convolutional Networks (Stacked-STGCN) for action segmentation, i.e., predicting and localizing a sequence of actions over long video…