3 papers
cs.CL2026
Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
Ping Wu, Haibo Tong, Feifei Zhao +7
Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refus…
cs.LG2026
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
Bing Han, Feifei Zhao, Dongcheng Zhao +4
While fine-tuning services drive the rapid expansion of task capabilities in large language models (LLMs), they are often accompanied by the degradation and reorganization of safet…
cs.CL2026
C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models
Ping Wu, Guobin Shen, Dongcheng Zhao +6
Ensuring that Large Language Models (LLMs) align with mainstream human values and ethical norms is crucial for the safe and sustainable development of AI. Current value evaluation…