3 citations · 8 across the 6 of their papers we have counts for
6 papers · 1 filter
When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning
Yitong Guo, Xiaoyi Chen, Siyuan Zhang +2
Benign fine-tuning severely weakens the safety alignment of large language models (LLMs), so we study why refusal behavior is so fragile. While prior work often attributes this fai…
LLM-Enhanced Software Patch Localization
Jinhong Yu, Yi Chen, Di Tang +4
Open source software (OSS) is integral to modern product development, and any vulnerability within it potentially compromises numerous products. While developers strive to apply se…
The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
Xiaoyi Chen, Siyuan Tang, Rui Zhu +7
The rapid advancements of large language models (LLMs) have raised public concerns about the privacy leakage of personally identifiable information (PII) within their extensive tra…
MAWSEO: Adversarial Wiki Search Poisoning for Illicit Online Promotion
Zilong Lin, Zhengyi Li, Xiaojing Liao +2
As a prominent instance of vandalism edits, Wiki search poisoning for illicit promotion is a cybercrime in which the adversary aims at editing Wiki articles to promote illicit busi…
Understanding Impacts of Task Similarity on Backdoor Attack and Detection
Di Tang, Rui Zhu, XiaoFeng Wang +2
With extensive studies on backdoor attack and detection, still fundamental questions are left unanswered regarding the limits in the adversary's capability to attack and the defend…
Confidential Attestation: Efficient in-Enclave Verification of Privacy Policy Compliance
Weijie Liu, Wenhao Wang, Xiaofeng Wang +9
A trusted execution environment (TEE) such as Intel Software Guard Extension (SGX) runs a remote attestation to prove to a data owner the integrity of the initial state of an encla…