6 papers
PRSM: A Measure to Evaluate CLIP's Robustness Against Paraphrases
Udo Schlegel, Franziska Weeber, Jian Lan +1
Contrastive Language-Image Pre-training (CLIP) is a widely used multimodal model that aligns text and image representations through large-scale training. While it performs strongly…
Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
Jian Lan, Zhicheng Liu, Udo Schlegel +5
Large vision-language models (VLMs) achieve strong performance in Visual Question Answering but still rely heavily on supervised fine-tuning (SFT) with massive labeled datasets, wh…
Enhancing Robustness of Autoregressive Language Models against Orthographic Attacks via Pixel-based Approach
Han Yang, Jian Lan, Yihong Liu +2
Autoregressive language models are vulnerable to orthographic attacks, where input text is perturbed with characters from multilingual alphabets, leading to substantial performance…
AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
Gengyuan Zhang, Tanveer Hannan, Hermine Kleiner +6
An ideal vision-language agent serves as a bridge between the human users and their surrounding physical world in real-world applications like autonomous driving and embodied agent…
Does Machine Unlearning Truly Remove Knowledge?
Haokun Chen, Yueqi Zhang, Yuan Bi +9
In recent years, Large Language Models (LLMs) have achieved remarkable advancements, drawing significant attention from the research community. Their capabilities are largely attri…
Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA
Jian Lan, Diego Frassinelli, Barbara Plank
Large vision-language models frequently struggle to accurately predict responses provided by multiple human annotators, particularly when those responses exhibit human uncertainty.…