6 papers · 1 filter
3-Tracer: A Tri-level Temporal-Aware Framework for Audio Forgery Detection and Localization
Shuhan Xia, Xuannan Liu, Xing Cui +1
Recently, partial audio forgery has emerged as a new form of audio manipulation. Attackers selectively modify partial but semantically critical frames while preserving the overall…
SpineBench: Benchmarking Multimodal LLMs for Spinal Pathology Analysis
Chenghanyu Zhang, Zekun Li, Peipei Li +5
With the increasing integration of Multimodal Large Language Models (MLLMs) into the medical field, comprehensive evaluation of their performance in various medical domains becomes…
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
Xuannan Liu, Zekun Li, Zheqi He +6
The increasing deployment of Large Vision-Language Models (LVLMs) raises safety concerns under potential malicious inputs. However, existing multimodal safety evaluations primarily…
ID-Cloak: Crafting Identity-Specific Cloaks Against Personalized Text-to-Image Generation
Qianrui Teng, Xing Cui, Xuannan Liu +4
Personalized text-to-image models allow users to generate images of new concepts from several reference photos, thereby leading to critical concerns regarding civil privacy. Althou…
Survey on AI-Generated Media Detection: From Non-MLLM to MLLM
Yueying Zou, Peipei Li, Zekun Li +5
The proliferation of AI-generated media poses significant challenges to information authenticity and social trust, making reliable detection methods highly demanded. Methods for de…
Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
Xuannan Liu, Xing Cui, Peipei Li +6
The rapid evolution of multimodal foundation models has led to significant advancements in cross-modal understanding and generation across diverse modalities, including text, image…