3 citations · 9 across the 20 of their papers we have counts for
6 papers · 1 filter
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
Ashish Seth, Dinesh Manocha, Chirag Agarwal
Large Vision-Language Models (LVLMs) have demonstrated remarkable performance in complex multimodal tasks. However, these models still suffer from hallucinations, particularly when…
Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar +4
Large Vision-Language Models (LVLMs) often produce responses that misalign with factual information, a phenomenon known as hallucinations. While hallucinations are well-studied, th…
Do Vision-Language Models Understand Compound Nouns?
Sonal Kumar, Sreyan Ghosh, S Sakshi +2
Open-vocabulary vision-language models (VLMs) like CLIP, trained using contrastive loss, have emerged as a promising new paradigm for text-to-image retrieval. However, do VLMs unde…
ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations
Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar +4
Neural image classifiers can often learn to make predictions by overly relying on non-predictive features that are spuriously correlated with the class labels in the training data.…
LoLep: Single-View View Synthesis with Locally-Learned Planes and Self-Attention Occlusion Inference
Cong Wang, Yu-Ping Wang, Dinesh Manocha
We propose a novel method, LoLep, which regresses Locally-Learned planes from a single RGB image to represent scenes accurately, thus generating better novel views. Without the dep…
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
Xijun Wang, Ruiqi Xian, Tianrui Guan +2
We present a new learning approach, Soft Conditional Prompt Learning (SCP), which leverages the strengths of prompt learning for aerial video action recognition. Our approach is de…