6 papers
Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization
Zhixin Lin, Jungang Li, Dongliang Xu +5
Mobile GUI agents powered by Multimodal Large Language Models (MLLMs) can execute complex tasks on mobile devices. Despite this progress, most existing systems still optimize task…
ACT Now: Preempting LVLM Hallucinations via Adaptive Context Integration
Bei Yan, Yuecong Min, Jie Zhang +2
Large Vision-Language Models (LVLMs) frequently suffer from severe hallucination issues. Existing mitigation strategies predominantly rely on isolated, single-step states to enhanc…
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
Junqi Yang, Yuecong Min, Jie Zhang +2
Despite rapid progress, Video Large Language Models (Video-LLMs) remain unreliable due to hallucinations, which are outputs that contradict either video evidence (faithfulness) or…
A Survey of Multimodal Hallucination Evaluation and Detection
Zhiyuan Chen, Yuecong Min, Jie Zhang +4
Multi-modal Large Language Models (MLLMs) have emerged as a powerful paradigm for integrating visual and textual information, supporting a wide range of multi-modal tasks. However,…
T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models
Changzhen Li, Yuecong Min, Jie Zhang +3
The rapid evolution of Text-to-Video (T2V) diffusion models has driven remarkable advancements in generating high-quality, temporally coherent videos from natural language descript…
SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
Bei Yan, Zhiyuan Chen, Yuecong Min +4
Despite rapid advances, Large Vision-Language Models (LVLMs) still suffer from hallucinations, i.e., generating content inconsistent with input or established world knowledge, whic…