collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Leveraging Latent Visual Reasoning in Silence

Dongyao Zhu, Zhen Wang, Xi Xiao +7

Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of th…

cs.CV2026

WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments

Haoren Zhao, Tianyi Chen, Zhen Wang

Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interfaces. This paper identifies…

cs.CV2026

Exposing and Mitigating Temporal Attack in Deepfake Video Detection

Zheyuan Gu, Minghao Shao, Zhen Wang +4

While spatiotemporal deepfake detectors achieve high AUC, our experiments reveal their susceptibility to evasion attacks. These models tend to overfit on fragile temporal spectrum…

cs.CV2025

Evaluating Adversarial Protections for Diffusion Personalization: A Comprehensive Study

Kai Ye, Tianyi Chen, Zhen Wang

With the increasing adoption of diffusion models for image generation and personalization, concerns regarding privacy breaches and content misuse have become more pressing. In this…

cs.CV2025

On the Robustness of GUI Grounding Models Against Image Attacks

Haoren Zhao, Tianyi Chen, Zhen Wang

Graphical User Interface (GUI) grounding models are crucial for enabling intelligent agents to understand and interact with complex visual interfaces. However, these models face si…