3 papers
cs.CV2026
Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
Ruoxi Cheng, Haoxuan Ma, Zhengfei Hai +6
Large Vision-Language Models (LVLMs) have advanced multimodal understanding, yet their reliability is limited by hallucination, where generated content conflicts with visual facts.…
cs.AI2025
ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection
Xin Zhang, Jiaming Chu, Jian Zhao +5
Deepfake detection is a critical task in identifying manipulated multimedia content. In real-world scenarios, deepfake content can manifest across multiple modalities, including au…
cs.LG2025
METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark
Xu Yang, Qi Zhang, Shuming Jiang +4
With the rapid advancement of generative AI, synthetic content across images, videos, and audio has become increasingly realistic, amplifying the risk of misinformation. Existing d…