activity
20202024
most citedExploiting BERT For Multimodal Target Sentiment Classification Through Input Space Translation

190 citations · 233 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2024

Consistency and Uncertainty: Identifying Unreliable Responses From Black-Box Vision-Language Models for Selective Visual Question Answering

Zaid Khan, Yun Fu

The goal of selective prediction is to allow an a model to abstain when it may not be able to deliver a reliable prediction, which is important in safety-critical contexts. Existin…

cs.CV2024

Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement

Zaid Khan, Vijay Kumar BG, Samuel Schulter +2

Visual program synthesis is a promising approach to exploit the reasoning abilities of large language models for compositional computer vision tasks. Previous work has used few-sho…

cs.CV2023

Exploring Question Decomposition for Zero-Shot VQA

Zaid Khan, Vijay Kumar BG, Samuel Schulter +2

Visual question answering (VQA) has traditionally been treated as a single-step task where each question receives the same amount of effort, unlike natural human question-answering…

cs.CV2023

Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!

Zaid Khan, Vijay Kumar BG, Samuel Schulter +3

Finetuning a large vision language model (VLM) on a target dataset after large scale pretraining is a dominant paradigm in visual question answering (VQA). Datasets for specialized…

cs.CV20232 cited

Contrastive Alignment of Vision to Language Through Parameter-Efficient Transfer Learning

Zaid Khan, Yun Fu

Contrastive vision-language models (e.g. CLIP) are typically created by updating all the parameters of a vision model and language model through contrastive training. Can such mode…

cs.CV202141 cited

One Label, One Billion Faces: Usage and Consistency of Racial Categories in Computer Vision

Zaid Khan, Yun Fu

Computer vision is widely deployed, has highly visible, society altering applications, and documented problems with bias and representation. Datasets are critical for benchmarking…