9 papers
AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models
Zhiwei Li, Jiacheng Xue, Weining Wang +4
Vision-language models like CLIP have demonstrated remarkable zero-shot transfer capabilities. However, their susceptibility to imperceptible adversarial perturbations remains a cr…
Towards High Fidelity Face Swapping: A Comprehensive Survey and New Benchmark
Qi Li, Weining Wang, Shuangjun Du +5
Face swapping has witnessed significant progress in recent years, largely driven by advances in deep generative models such as GANs and diffusion models.Despite these advances, exi…
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
Qianshan Wei, Yishan Yang, Siyi Wang +12
Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Exp…
3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
Xinyang Song, Libin Wang, Weining Wang +6
Recent image generation approaches often address subject, style, and structure-driven conditioning in isolation, leading to feature entanglement and limited task transferability. I…
PhysCorr: Dual-Reward DPO for Physics-Constrained Text-to-Video Generation with Automated Preference Selection
Peiyao Wang, Weining Wang, Qi Li
Recent advances in text-to-video generation have achieved impressive perceptual quality, yet generated content often violates fundamental principles of physical plausibility - mani…
Learning Knowledge-based Prompts for Robust 3D Mask Presentation Attack Detection
Fangling Jiang, Qi Li, Bing Liu +4
3D mask presentation attack detection is crucial for protecting face recognition systems against the rising threat of 3D mask attacks. While most existing methods utilize multimoda…