2 citations · 2 across the 9 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
Ping Wu, Haibo Tong, Feifei Zhao +7
Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refus…
cs.CL2026
C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models
Ping Wu, Guobin Shen, Dongcheng Zhao +6
Ensuring that Large Language Models (LLMs) align with mainstream human values and ethical norms is crucial for the safe and sustainable development of AI. Current value evaluation…
cs.CL2025
DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values
Yao Liang, Dongcheng Zhao, Feifei Zhao +4
Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks often overlook cultural and demographic vari…