2 papers
cs.CR2025
SDD: Self-Degraded Defense against Malicious Fine-tuning
Zixuan Chen, Weikai Lu, Xin Lin +1
Open-source Large Language Models (LLMs) often employ safety alignment methods to resist harmful instructions. However, recent research shows that maliciously fine-tuning these LLM…
cs.CL2024
Zero-shot Explainable Mental Health Analysis on Social Media by Incorporating Mental Scales
Wenyu Li, Yinuo Zhu, Xin Lin +3
Traditional discriminative approaches in mental health analysis are known for their strong capacity but lack interpretability and demand large-scale annotated data. The generative…