3 papers
cs.CV2026
Attribution-Guided Multimodal Deepfake Detection via Cross-Modal Forensic Fingerprints
Wasim Ahmad, Wei Zhang, Xuerui Mao
Audio-visual deepfakes have reached a level of realism that makes perceptual detection unreliable, threatening media integrity and biometric security. While multimodal detection ha…
cs.CV2025
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
Wei Zhang, Miaoxin Cai, Yaqian Ning +6
Recent advances in natural-domain multi-modal large language models (MLLMs) have demonstrated effective spatial reasoning through visual and textual prompting. However, their direc…
cs.CV2025
RS-TinyNet: Stage-wise Feature Fusion Network for Detecting Tiny Objects in Remote Sensing Images
Xiaozheng Jiang, Wei Zhang, Xuerui Mao
Detecting tiny objects in remote sensing (RS) imagery has been a long-standing challenge due to their extremely limited spatial information, weak feature representations, and dense…