activity
20222024
most citedExploiting Unlabeled Data with Vision and Language Models for Object Detection

2 citations · 3 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV2024

Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement

Zaid Khan, Vijay Kumar BG, Samuel Schulter +2

Visual program synthesis is a promising approach to exploit the reasoning abilities of large language models for compositional computer vision tasks. Previous work has used few-sho…

cs.CV2024

AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving

Mingfu Liang, Jong-Chyi Su, Samuel Schulter +4

Autonomous vehicle (AV) systems rely on robust perception models as a cornerstone of safety assurance. However, objects encountered on the road exhibit a long-tailed distribution,…

cs.CV2024

Generating Enhanced Negatives for Training Language-Based Object Detectors

Shiyu Zhao, Long Zhao, Vijay Kumar B. G +4

The recent progress in language-based open-vocabulary object detection can be largely attributed to finding better ways of leveraging large-scale data with free-form text annotatio…

cs.CV2023

Exploring Question Decomposition for Zero-Shot VQA

Zaid Khan, Vijay Kumar BG, Samuel Schulter +2

Visual question answering (VQA) has traditionally been treated as a single-step task where each question receives the same amount of effort, unlike natural human question-answering…

cs.CV2023

Efficient Controllable Multi-Task Architectures

Abhishek Aich, Samuel Schulter, Amit K. Roy-Chowdhury +2

We aim to train a multi-task model such that users can adjust the desired compute budget and relative importance of task performances after deployment, without retraining. This ena…

cs.CV2023

Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!

Zaid Khan, Vijay Kumar BG, Samuel Schulter +3

Finetuning a large vision language model (VLM) on a target dataset after large scale pretraining is a dominant paradigm in visual question answering (VQA). Datasets for specialized…