2 citations · 3 across the 21 of their papers we have counts for
33 papers · 1 filter
Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model
Sibo Wang, Jie Zhang, Shiguang Shan +2
While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs remain vulnerable to adversarial attacks…
EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy
Xuanyu Ge, Zhongqi Wang, Jie Zhang +2
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to backdoor attacks. Existing defense methods predom…
Component-Based Out-of-Distribution Detection
Wenrui Liu, Hong Chang, Ruibing Hou +2
Out-of-Distribution (OOD) detection requires sensitivity to subtle shifts without overreacting to natural In-Distribution (ID) diversity. However, from the viewpoint of detection g…
EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation
Ruibing Hou, Mingyue Zhou, Yuwei Gui +5
Faithfully modeling human behavior in dynamic environments is a foundational challenge for embodied intelligence. While conditional motion synthesis has achieved significant advanc…
ACT Now: Preempting LVLM Hallucinations via Adaptive Context Integration
Bei Yan, Yuecong Min, Jie Zhang +2
Large Vision-Language Models (LVLMs) frequently suffer from severe hallucination issues. Existing mitigation strategies predominantly rely on isolated, single-step states to enhanc…
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
Junqi Yang, Yuecong Min, Jie Zhang +2
Despite rapid progress, Video Large Language Models (Video-LLMs) remain unreliable due to hallucinations, which are outputs that contradict either video evidence (faithfulness) or…