6 papers
ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents
Jianan Ma, Xiaohu Du, Ruixiao Lin +9
As autonomous agents (e.g., OpenClaw) increasingly operate with deep system-level privileges to execute complex tasks, they introduce severe, unmitigated security risks. Existing L…
Leveraging Latent Visual Reasoning in Silence
Dongyao Zhu, Zhen Wang, Xi Xiao +7
Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of th…
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
Haoren Zhao, Tianyi Chen, Zhen Wang
Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interfaces. This paper identifies…
Exposing and Mitigating Temporal Attack in Deepfake Video Detection
Zheyuan Gu, Minghao Shao, Zhen Wang +4
While spatiotemporal deepfake detectors achieve high AUC, our experiments reveal their susceptibility to evasion attacks. These models tend to overfit on fragile temporal spectrum…
Evaluating Adversarial Protections for Diffusion Personalization: A Comprehensive Study
Kai Ye, Tianyi Chen, Zhen Wang
With the increasing adoption of diffusion models for image generation and personalization, concerns regarding privacy breaches and content misuse have become more pressing. In this…
On the Robustness of GUI Grounding Models Against Image Attacks
Haoren Zhao, Tianyi Chen, Zhen Wang
Graphical User Interface (GUI) grounding models are crucial for enabling intelligent agents to understand and interact with complex visual interfaces. However, these models face si…