7 citations · 15 across the 7 of their papers we have counts for
7 papers · 1 filter
Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs
Yu Cheng, Arushi Goel, Hakan Bilen
Tool-augmented multimodal reasoning integrates external tools (e.g., object detection, depth estimation) into multimodal large language models (MLLMs) to address perceptual bottlen…
Visually Interpretable Subtask Reasoning for Visual Question Answering
Yu Cheng, Arushi Goel, Hakan Bilen
Answering complex visual questions like `Which red furniture can be used for sitting?' requires multi-step reasoning, including object recognition, attribute filtering, and relatio…
Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories
Thomas Mensink, Jasper Uijlings, Lluis Castrejon +6
We propose Encyclopedic-VQA, a large scale visual question answering (VQA) dataset featuring visual questions about detailed properties of fine-grained categories and instances. It…
WiCV 2021: The Eighth Women In Computer Vision Workshop
Arushi Goel, Niveditha Kalavakonda, Nour Karessli +5
In this paper, we present the details of Women in Computer Vision Workshop - WiCV 2021, organized alongside the virtual CVPR 2021. It provides a voice to a minority (female) group…
PARS: Pseudo-Label Aware Robust Sample Selection for Learning with Noisy Labels
Arushi Goel, Yunlong Jiao, Jordan Massiah
Acquiring accurate labels on large-scale datasets is both time consuming and expensive. To reduce the dependency of deep learning models on learning from clean labeled data, severa…
Cross-Domain Image Classification through Neural-Style Transfer Data Augmentation
Yijie Xu, Arushi Goel
In particular, the lack of sufficient amounts of domain-specific data can reduce the accuracy of a classifier. In this paper, we explore the effects of style transfer-based data tr…