collaborators

6 papers

cs.CR2026

LoopTrap: Termination Poisoning Attacks on LLM Agents

Huiyu Xu, Zhibo Wang, Wenhui Zhang +4

Modern LLM agents solve complex tasks by operating in iterative execution loops, where they repeatedly reason, act, and self-evaluate progress to determine when a task is complete.…

cs.CR2025

DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization

Xinzhe Huang, Kedong Xiu, Tianhang Zheng +5

Recent research has focused on exploring the vulnerabilities of Large Language Models (LLMs), aiming to elicit harmful and/or sensitive content from LLMs. However, due to the insuf…

cs.CR2025

Combating Concept Drift with Explanatory Detection and Adaptation for Android Malware Classification

Yiling He, Junchi Lei, Zhan Qin +2

Machine learning-based Android malware classifiers achieve high accuracy in stationary environments but struggle with concept drift. The rapid evolution of malware, especially with…

cs.CR2025

JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit

Zeqing He, Zhibo Wang, Zhixuan Chu +4

Despite the outstanding performance of Large language Models (LLMs) in diverse tasks, they are vulnerable to jailbreak attacks, wherein adversarial prompts are crafted to bypass th…

cs.CR2025

Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models

Yu He, Boheng Li, Liu Liu +6

Membership Inference Attacks (MIAs) aim to predict whether a data sample belongs to the model's training set or not. Although prior research has extensively explored MIAs in Large…

cs.CR2025

Membership Inference Attacks Against Vision-Language Models

Yuke Hu, Zheng Li, Zhihao Liu +4

Vision-Language Models (VLMs), built on pre-trained vision encoders and large language models (LLMs), have shown exceptional multi-modal understanding and dialog capabilities, posi…