activity
20242026
most citedJailbreak Attacks and Defenses Against Large Language Models: A Survey

7 citations · 8 across the 5 of their papers we have counts for

collaborators

9 papers

cs.CL2026

Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis

Da Song, Yuheng Huang, Boqi Chen +4

The integration of large language models (LLMs) into autonomous agents has enabled complex tool use, yet in high-stakes domains, these systems must strictly adhere to regulatory st…

cs.CR2025

LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models

Delong Ran, Xinlei He, Tianshuo Cong +3

Language Models (LMs) typically adhere to a "pre-training and fine-tuning" paradigm, where a universal pre-trained model can be fine-tuned to cater to various specialized domains.…

cs.CR2025

Watermarking LLM-Generated Datasets in Downstream Tasks

Yugeng Liu, Tianshuo Cong, Michael Backes +2

Large Language Models (LLMs) have experienced rapid advancements, with applications spanning a wide range of fields, including sentiment classification, review generation, and ques…

cs.CV2025

Can VLMs Detect and Localize Fine-Grained AI-Edited Images?

Zhen Sun, Ziyi Zhang, Zeren Luo +10

Fine-grained detection and localization of localized image edits is crucial for assessing content authenticity, especially as modern diffusion models and image editors can produce…

cs.CR20251 cited

SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning

Heyi Zhang, Yule Liu, Xinlei He +3

Federated learning (FL) enables collaborative model training while preserving data privacy, but its decentralized nature exposes it to client-side data poisoning attacks (DPAs) and…

cs.CR2025

Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models

Sibo Yi, Tianshuo Cong, Xinlei He +2

Small language models (SLMs) have become increasingly prominent in the deployment on edge devices due to their high efficiency and low computational cost. While researchers continu…