3 citations · 6 across the 8 of their papers we have counts for
8 papers
T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models
Changzhen Li, Yuecong Min, Jie Zhang +3
The rapid evolution of Text-to-Video (T2V) diffusion models has driven remarkable advancements in generating high-quality, temporally coherent videos from natural language descript…
REVAL: A Comprehension Evaluation on Reliability and Values of Large Vision-Language Models
Jie Zhang, Zheng Yuan, Zhongqi Wang +6
The rapid evolution of Large Vision-Language Models (LVLMs) has highlighted the necessity for comprehensive evaluation frameworks that assess these models across diverse dimensions…
Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs
Jie Zhang, Zhongqi Wang, Mengqi Lei +4
Currently many benchmarks have been proposed to evaluate the perception ability of the Large Vision-Language Models (LVLMs). However, most benchmarks conduct questions by selecting…
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
Bei Yan, Jie Zhang, Zheng Yuan +2
Despite the outstanding performance in multimodal tasks, Large Vision-Language Models (LVLMs) have been plagued by the issue of hallucination, i.e., generating content that is inco…
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
Sibo Wang, Xiangkui Cao, Jie Zhang +4
The emergence of Large Vision-Language Models (LVLMs) marks significant strides towards achieving general artificial intelligence. However, these advancements are accompanied by co…
Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness
Sibo Wang, Jie Zhang, Zheng Yuan +1
Large-scale pre-trained vision-language models like CLIP have demonstrated impressive performance across various tasks, and exhibit remarkable zero-shot generalization capability,…