3 papers
cs.LG2026
Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs
Xin Zhang, Qiqi Tao, Jiawei Du +2
Continuous latent-space reasoning offers a compact alternative to textual chain-of-thought for multimodal models, enabling high-dimensional visual evidence to be integrated without…
cs.CV2025
TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models
Zhifang Zhang, Qiqi Tao, Jiaqi Lv +3
Large vision-language models (LVLMs) have achieved impressive performance across a wide range of vision-language tasks, while they remain vulnerable to backdoor attacks. Existing b…
cs.CV2024
Evolving from Single-modal to Multi-modal Facial Deepfake Detection: Progress and Challenges
Ping Liu, Qiqi Tao, Joey Tianyi Zhou
As synthetic media, including video, audio, and text, become increasingly indistinguishable from real content, the risks of misinformation, identity fraud, and social manipulation…