6 papers
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
Sibo Wang, Xiangkui Cao, Jie Zhang +4
The emergence of Large Vision-Language Models (LVLMs) marks significant strides towards achieving general artificial intelligence. However, these advancements are accompanied by co…
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
Bei Yan, Jie Zhang, Zheng Yuan +2
Despite the outstanding performance in multimodal tasks, Large Vision-Language Models (LVLMs) have been plagued by the issue of hallucination, i.e., generating content that is inco…
T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models
Changzhen Li, Yuecong Min, Jie Zhang +3
The rapid evolution of Text-to-Video (T2V) diffusion models has driven remarkable advancements in generating high-quality, temporally coherent videos from natural language descript…
FullLoRA: Efficiently Boosting the Robustness of Pretrained Vision Transformers
Zheng Yuan, Jie Zhang, Shiguang Shan +1
In recent years, the Vision Transformer (ViT) model has gradually become mainstream in various computer vision tasks, and the robustness of the model has received increasing attent…
REVAL: A Comprehension Evaluation on Reliability and Values of Large Vision-Language Models
Jie Zhang, Zheng Yuan, Zhongqi Wang +6
The rapid evolution of Large Vision-Language Models (LVLMs) has highlighted the necessity for comprehensive evaluation frameworks that assess these models across diverse dimensions…
Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs
Jie Zhang, Zhongqi Wang, Mengqi Lei +4
Currently many benchmarks have been proposed to evaluate the perception ability of the Large Vision-Language Models (LVLMs). However, most benchmarks conduct questions by selecting…