collaborators

6 papers

cs.CR2026

ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents

Jianan Ma, Xiaohu Du, Ruixiao Lin +9

As autonomous agents (e.g., OpenClaw) increasingly operate with deep system-level privileges to execute complex tasks, they introduce severe, unmitigated security risks. Existing L…

cs.CV2026

Leveraging Latent Visual Reasoning in Silence

Dongyao Zhu, Zhen Wang, Xi Xiao +7

Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of th…

cs.CV2026

WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments

Haoren Zhao, Tianyi Chen, Zhen Wang

Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interfaces. This paper identifies…

cs.CV2026

Exposing and Mitigating Temporal Attack in Deepfake Video Detection

Zheyuan Gu, Minghao Shao, Zhen Wang +4

While spatiotemporal deepfake detectors achieve high AUC, our experiments reveal their susceptibility to evasion attacks. These models tend to overfit on fragile temporal spectrum…

cs.CV2025

Evaluating Adversarial Protections for Diffusion Personalization: A Comprehensive Study

Kai Ye, Tianyi Chen, Zhen Wang

With the increasing adoption of diffusion models for image generation and personalization, concerns regarding privacy breaches and content misuse have become more pressing. In this…

cs.CV2025

On the Robustness of GUI Grounding Models Against Image Attacks

Haoren Zhao, Tianyi Chen, Zhen Wang

Graphical User Interface (GUI) grounding models are crucial for enabling intelligent agents to understand and interact with complex visual interfaces. However, these models face si…