4 papers · 1 filter
Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection
Tianxiao Li, Zhenglin Huang, Haiquan Wen +10
Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are constrained by their single-…
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
Xinmiao Huang, Qisong He, Zhenglin Huang +5
Spatial reasoning ability is crucial for Vision Language Models (VLMs) to support real-world applications in diverse domains including robotics, augmented reality, and autonomous n…
TAIJI: Textual Anchoring for Immunizing Jailbreak Images in Vision Language Models
Xiangyu Yin, Yi Qi, Jinwei Hu +5
Vision Language Models (VLMs) have demonstrated impressive inference capabilities, but remain vulnerable to jailbreak attacks that can induce harmful or unethical responses. Existi…
CeTAD: Towards Certified Toxicity-Aware Distance in Vision Language Models
Xiangyu Yin, Jiaxu Liu, Zhen Chen +4
Recent advances in large vision-language models (VLMs) have demonstrated remarkable success across a wide range of visual understanding tasks. However, the robustness of these mode…