5 papers
UGID: Unified Graph Isomorphism for Debiasing Large Language Models
Zikang Ding, Junchi Yao, Junhao Li +4
Large language models (LLMs) exhibit pronounced social biases. Output-level or data-optimization--based debiasing methods cannot fully resolve these biases, and many prior works ha…
Functional Subspace Watermarking for Large Language Models
Zikang Ding, Junhao Li, Suling Wu +3
Model watermarking utilizes internal representations to protect the ownership of large language models (LLMs). However, these features inevitably undergo complex distortions during…
FaithSteer-BENCH: A Deployment-Aligned Stress-Testing Benchmark for Inference-Time Steering
Zikang Ding, Qiying Hu, Yi Zhang +4
Inference-time steering is widely regarded as a lightweight and parameter-free mechanism for controlling large language model (LLM) behavior, and prior work has often suggested tha…
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
Zikang Ding, Haomiao Yang, Meng Hao +6
Backdoor attacks against pre-trained models (PTMs) have traditionally operated under an ``immediacy assumption,'' where malicious behavior manifests instantly upon trigger occurren…
The Gradient Puppeteer: Adversarial Domination in Gradient Leakage Attacks through Model Poisoning
Kunlan Xiang, Haomiao Yang, Meng Hao +5
In Federated Learning (FL), clients share gradients with a central server while keeping their data local. However, malicious servers could deliberately manipulate the models to rec…