Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models
Yuhang Wang, Yanxu Zhu, Jiaming Zhang +2
Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We study whether a reasoning model can synth…
cs.CL2023
You talk what you read: Understanding News Comment Behavior by Dispositional and Situational Attribution
Yuhang Wang, Yuxiang Zhang, Dongyuan Lu +1
Many news comment mining studies are based on the assumption that comment is explicitly linked to the corresponding news. In this paper, we observed that users' comments are also h…