Showing cs.CRShow all
2 papers · 1 filter
cs.CR2024★ 2 cited
Safety Layers in Aligned Large Language Models: The Key to LLM Security
Shen Li, Liuyi Yao, Lan Zhang +1
Aligned LLMs are secure, capable of recognizing and refusing to answer malicious questions. However, the role of internal parameters in maintaining such security is not well unders…
cs.CR2024
Double-I Watermark: Protecting Model Copyright for LLM Fine-tuning
Shen Li, Liuyi Yao, Jinyang Gao +2
To support various applications, a prevalent and efficient approach for business owners is leveraging their valuable datasets to fine-tune a pre-trained LLM through the API provide…