3 papers
cs.CR2026
DNF: Dual-Layer Nested Fingerprinting for Large Language Model Intellectual Property Protection
Zhenhua Xu, Yiran Zhao, Mengting Zhong +4
The rapid growth of large language models raises pressing concerns about intellectual property protection under black-box deployment. Existing backdoor-based fingerprints either re…
cs.CR2025
Black-Box Guardrail Reverse-engineering Attack
Hongwei Yao, Yun Xia, Shuo Shao +3
Large language models (LLMs) increasingly employ guardrails to enforce ethical, legal, and application-specific constraints on their outputs. While effective at mitigating harmful…
cs.LG2025
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
Kaiwen Duan, Hongwei Yao, Yufei Chen +4
Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning text-to-image (T2I) models with human preferences. However, RLHF's feedback mechanism also opens new pathw…