3 papers
cs.CL2025
Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models
Yuhang Wang, Yanxu Zhu, Jiaming Zhang +2
Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We study whether a reasoning model can synth…
cs.LG2024
Inference-Time Rule Eraser: Fair Recognition via Distilling and Removing Biased Rules
Yi Zhang, Dongyuan Lu, Jitao Sang
Machine learning models often make predictions based on biased features such as gender, race, and other social attributes, posing significant fairness risks, especially in societal…
cs.CL2023
You talk what you read: Understanding News Comment Behavior by Dispositional and Situational Attribution
Yuhang Wang, Yuxiang Zhang, Dongyuan Lu +1
Many news comment mining studies are based on the assumption that comment is explicitly linked to the corresponding news. In this paper, we observed that users' comments are also h…